Skip to main content

Benchmark suite map

Suites have different prerequisites: startup and the fx comparison need Hyperfine; the latter also needs an executable fx. A live harness comparison needs external CLIs and OpenRouter credentials; review its dry run before running. See the repository’s benchmarks/README.md and benchmarks/harness-comparison/README.md for setup and methodology. Generated summaries under benchmarks/results/ are not uniformly checked in; the figures below are historical local observations, not universal promises.

Kernel dispatch methodology

fanout_probe runs deterministic no-op calls at batch sizes 1, 10, and 100. It records T0 at dispatch entry, T1 when the first tool body begins, and T2 when the last tool body begins. The standard profile uses 2,000 iterations. --ci uses 5,000 so p99 remains a percentile on a noisy shared runner. --quick reduces iteration count and skips real-tool probes. --criterion additionally runs workspace Criterion benches and folds their estimates into the summary.

Historical kernel observations

These previously reported local values are not backed by a checked-in kernel summary in this checkout; rerun the suite for current host-specific results. Real write_file, read_file, and shell probes run through the same dispatcher, but remain informational because filesystem and process-spawn performance belong substantially to the host.

Cold-start fixtures

The startup runner uses Hyperfine against a generated private HOME, config directory, workspaces, sessions, and skills. It unsets provider keys and XDG_CONFIG_HOME, preventing a developer’s machine configuration from changing the path under test.
The benchmark hook: ORCA_BENCH=1 exits after run_mode builds the agent and immediately before the TUI claims the terminal. It awaits MCP initialization, performs no model call, and requires no TTY. ORCA_BENCH=ui measures the shorter pre-terminal construction path without launching background MCP connections; it excludes rendering and is not a full-integration-readiness measurement.

Historical startup observations

These are historical observations, not a fresh validation of the checkout against current gates; the displayed startup and new-session means do not demonstrate a pass under the strict ceiling. The resume fixture ends on complete user/assistant/tool triplets and is never rewritten. The session-creation case receives a separate workspace whose session directory is reset before each run. Those details prevent fixture repair or accumulated files from contaminating the measurement.

Memory retrieval and mutation overhead

The Criterion fixture creates 50,000 records: 5,000 global and 45,000 in one workspace. Exact and broad queries compare direct scoped FTS5 SQL with MemoryStore::search and complete MemoryExtension::prepare_context hydration. The mutation pair compares raw scoped delete with memory_manage/forget_no_approval.
Measurement boundary: The benchmark calls the memory tool directly. It excludes terminal rendering, provider latency, and time spent waiting for a user to approve memory_manage. No memory latency result is checked in, so local Criterion output is evidence for that machine and build only.

MCP metadata-search accuracy

The deterministic MCP corpus contains sanitized metadata for 18 tools and 38 labeled query-target cases from enabled AWS Docs and read-only GitHub servers. It contains no credentials, server commands, or full input schemas. The benchmark mirrors production tokenization, field weights, exact-name bonuses, singular/plural fallback, tie-breaking, and stable catalog order. Required target ranks are 33 at rank 1, two at rank 2, and three at rank 3; none are unmatched. The measured Pareto knee is a limit of 3: it reaches 100% target coverage at 1.71 returned results per query. Larger limits add no coverage in this fixture.
This corpus is a regression fixture, not a universal MCP workload. Underspecified queries are intentionally ambiguous; their weak top-1 score should not be presented as retrieval failure when the labeled target remains in the returned set.

Fake-subagent concurrency stress

This suite drives the real SubagentTool::call and inner Agent path with a fake model that returns one answer after a Tokio timer. It isolates allocation, scheduling, model-boundary, and completion behavior without provider latency, API cost, network failure, or rate limits. The initial local reference run used macOS arm64 with 16 logical CPUs and 64 GiB RAM; its raw files are generated and git-ignored, not included in this checkout. Per-call latency starts at each spawned task’s first poll, excluding admission wait. peak_active counts futures simultaneously inside the delayed fake-model body, not CPUs executing simultaneously. A process killed by OOM cannot emit a CSV failure row.
Not a capacity claim: 131,072 is the largest level completed and directly observed in that machine-specific run—not a hard limit. The boundary suite is intentionally excluded from CI and performance budgets.

Comparing Orcacode with fx

The comparison subtracts a same-session /usr/bin/true process floor, then groups commands by work performed. It does not falsely pair benchmark environment variables that stop at different boundaries.
Do not turn this into a false race: FX_BENCH=1 parses argv and exits without reading settings. ORCA_BENCH=1 performs Orcacode’s complete cold-start construction. Conversely, fx status and doctor probe auth and the system, work Orcacode’s startup does not perform. Quote platform, build modes, command boundaries, and the process floor with any comparison.
This fx comparison measures CLI commands, not live model-driven task performance. The separate benchmarks/harness-comparison/ suite compares Orcacode, Pi, Oh My Pi, and Claude Code using OpenRouter-routed model calls. Its results depend on model routing, installed CLIs, and credentials; see its README and the historical report under benchmarks/results/ before quoting results.

Budgets, CI, and result integrity

benchmarks/shared/check_budgets.py gates startup means, selected kernel p99 values, and core-tools p50 latencies. These numerical budgets are enforced on Linux, where CI runs; other platforms report informationally unless ORCA_BENCH_ENFORCE=1 is set. CI also runs MCP metric reports with deterministic assertions in the suite’s tests. Workflow defect counts fail its analyzer when that suite is run; its latency and scale are informational.
Basic and new-session startup means must be strictly below 3.27 ms; equality fails. This is the ORCA_BENCH=1 pre-terminal path in an isolated fixture with no external MCP servers, not time to first rendered screen. Hyperfine uses warmups, and no process-floor subtraction is applied. These tight ceilings can be sensitive to shared-runner load; other scenarios retain looser regression ceilings. Separately, ci/check-binary-size.sh requires the uncompressed release binary to be strictly below 6,700,000 bytes (6.7 decimal MB).