# Criterion Benchmarks This file records the Criterion-backed benchmark suite for `aria2-rust-pro`. ## Purpose This project needs a repeatable benchmark surface in addition to assertion-style regression tests. This suite converts the existing synthetic BT/RPC pressure patterns into Criterion benchmarks so later tranches can compare revisions and resource-limit tuning with the same workload shape. ## Bench Target - crate: `crates/aria2-rust-pro-tests` - bench: `rpc_pressure` ## Current Coverage The current suite measures nine deterministic scenarios: 1. `rpc_tell_status_pressure` - 64-task and 128-task BT-like dispatcher loads - repeated `aria2.tellStatus` calls under runtime tick churn 2. `rpc_mixed_pressure` - 96-task and 192-task BT-like dispatcher loads - mixed `aria2.tellStatus`, `aria2.tellActive`, and `aria2.tellGlobalStat` batches 3. `runtime_snapshot_pressure` - 64 / 128 / 256 download in-memory engine setups - repeated scheduler runs plus `runtime_instrumentation_snapshot()` capture 4. `rpc_speed_limit_pressure` - 32-task and 64-task BT-like dispatcher loads - active `changeGlobalOption` / `changeOption` speed caps - repeated `tellStatus` plus `tellGlobalStat` under clamped runtime speeds 5. `scheduler_backpressure_pressure` - 64-task and 128-task in-memory engine setups - mixed active/waiting/error groups - repeated scheduler runs plus runtime instrumentation under differing `disk-cache` budgets 6. `live_http_transfer_contention_pressure` - loopback live HTTP transfer via `ReqwestHttpConnector` and `ConnectorBackedDownloader` - 16 KiB segmented HTTP downloads that force bootstrap `206` plus follow-up range requests - comparison between a loose-cap runtime and a tight `max-overall-download-limit` runtime so execution-layer throttling is measured instead of only inferred from RPC snapshots 7. `live_http_multi_download_contention_pressure` - three concurrent loopback live HTTP downloads - each download still uses the existing CLI/runtime wiring, but all three contend against one local segment server at once - compares loose-cap vs tight-cap runtime configuration under aggregate multi-download contention 8. `rpc_shared_runtime_fairness_pressure` - three-way and four-way BT-like dispatcher loads inside one runtime - one constrained gid completes, then the remaining active gids are expected to inherit a larger share of the global speed cap - verifies same-runtime cross-download rebalancing rather than only per-download segment throttling or separate parallel invocations 9. `live_http_shared_runtime_multi_download_pressure` - three loopback live HTTP downloads registered into one invocation/runtime - uses the shared CLI runtime path rather than separate parallel invocations - compares loose-cap vs tight-cap runtime configuration for same-runtime live transfer behavior - now also includes a six-download `tight_cap_6way` variant so shared-runtime cap behavior is measured beyond the original 3-download / 16 KiB shape - now also includes `cache_pressure_6way_256k`, which raises per-download transfer size while constraining `disk-cache` so shared-runtime cache pressure is represented in the same benchmark family ## Commands Canonical local report refresh: ```powershell rtk cargo run --manifest-path .\xtask\Cargo.toml -- perf collect-local-comparison pwsh ./scripts/perf/collect_local_comparison.ps1 ``` Compile-only verification: ```powershell rtk cargo bench --manifest-path .\Cargo.toml -p aria2-rust-pro-tests --bench rpc_pressure --no-run ``` Short local run: ```powershell rtk cargo bench --manifest-path .\Cargo.toml -p aria2-rust-pro-tests --bench rpc_pressure -- --sample-size 10 --measurement-time 0.05 --warm-up-time 0.05 ``` ## First Short Run Snapshot Short local run on this workspace produced the following indicative timings: - `rpc_tell_status_pressure/tell_status/64` - `1.3318 ms .. 1.3459 ms` - throughput `190.21 Kelem/s .. 192.22 Kelem/s` - `rpc_tell_status_pressure/tell_status/128` - `2.8581 ms .. 2.8726 ms` - throughput `178.24 Kelem/s .. 179.14 Kelem/s` - `rpc_mixed_pressure/mixed_rpc/96` - `59.918 us .. 64.561 us` - throughput `278.81 Kelem/s .. 300.41 Kelem/s` - `rpc_mixed_pressure/mixed_rpc/192` - `78.663 us .. 83.252 us` - throughput `216.21 Kelem/s .. 228.82 Kelem/s` - `runtime_snapshot_pressure/runtime_snapshot/64` - `8.3619 us .. 9.0689 us` - `runtime_snapshot_pressure/runtime_snapshot/128` - `15.452 us .. 17.835 us` - `runtime_snapshot_pressure/runtime_snapshot/256` - `31.072 us .. 33.087 us` - `rpc_speed_limit_pressure/speed_limit/32` - `64.896 us .. 66.191 us` - throughput `181.29 Kelem/s .. 184.91 Kelem/s` - `rpc_speed_limit_pressure/speed_limit/64` - `90.733 us .. 93.649 us` - throughput `128.14 Kelem/s .. 132.26 Kelem/s` - `scheduler_backpressure_pressure/backpressure/64` - `10.917 us .. 12.478 us` - throughput `5.1290 Melem/s .. 5.8625 Melem/s` - `scheduler_backpressure_pressure/backpressure/128` - `19.792 us .. 21.039 us` - throughput `6.0839 Melem/s .. 6.4673 Melem/s` - `live_http_transfer_contention_pressure/live_http_transfer/loose_cap` - `38.640 ms .. 39.143 ms` - throughput `408.76 KiB/s .. 414.07 KiB/s` - `live_http_transfer_contention_pressure/live_http_transfer/tight_cap` - `41.793 ms .. 48.988 ms` - throughput `326.61 KiB/s .. 382.83 KiB/s` These are not cross-machine release numbers. They are the first local benchmark local benchmark anchor so later tuning and regressions can be compared against a stable local benchmark driver. The new live HTTP group is still loopback-local rather than internet-realistic, but it now exercises the segmented execution path with real socket I/O instead of only scheduler-facing synthetic churn. - `live_http_multi_download_contention_pressure/multi_live_http_transfer/loose_cap` - `105.63 ms .. 106.46 ms` - throughput `450.86 KiB/s .. 454.43 KiB/s` - `live_http_multi_download_contention_pressure/multi_live_http_transfer/tight_cap` - `105.13 ms .. 105.52 ms` - throughput `454.90 KiB/s .. 456.56 KiB/s` - `rpc_shared_runtime_fairness_pressure/shared_runtime_fairness/three_way_rebalance` - `70.687 us .. 72.082 us` - throughput `166.48 Kelem/s .. 169.76 Kelem/s` - `rpc_shared_runtime_fairness_pressure/shared_runtime_fairness/four_way_rebalance` - `95.853 us .. 96.893 us` - throughput `165.13 Kelem/s .. 166.92 Kelem/s` - `live_http_shared_runtime_multi_download_pressure/shared_runtime_live_http_transfer/loose_cap` - `107.57 ms .. 108.23 ms` - throughput `443.49 KiB/s .. 446.23 KiB/s` - `live_http_shared_runtime_multi_download_pressure/shared_runtime_live_http_transfer/tight_cap` - `118.77 ms .. 120.35 ms` - throughput `398.83 KiB/s .. 404.14 KiB/s` - `live_http_shared_runtime_multi_download_pressure/shared_runtime_live_http_transfer/tight_cap_6way` - `234.36 ms .. 236.30 ms` - throughput `406.26 KiB/s .. 409.63 KiB/s` - `live_http_shared_runtime_multi_download_pressure/shared_runtime_live_http_transfer/cache_pressure_6way_256k` - `96.98 ms .. 99.12 ms` - throughput `15.133 MiB/s .. 15.466 MiB/s` ## Intended Next Extension This suite now mixes synthetic pressure and local loopback transfer. The next tranche should attach: - comparisons against the C++ aria2 reference and current Pro Core - further larger or more realistic same-runtime live transfer workloads beyond the new 6-download tight-cap and cache-pressure anchors - stronger disk-cache and backpressure scenarios with larger transfer/runtime contention - original C++ aria2 / current Pro Core comparison notes - memory and file-descriptor measurements recorded beside benchmark output