Benchmarks for the code-context MCP server — run on demand, never on a schedule.
Every published run, newest first, across all kinds. History is append-only: runs are never rewritten or deleted.
This page presents pre-registration v2 (registered 2026-10-09). The
authoritative text is PREREGISTRATION.md; v1 is frozen as
PREREGISTRATION-v1.md. The version and SHA-256 are recorded in every
results row and shown for every run. Changing the method after the first registered run
requires a new version with date and reason, and results are reported against the
version in force when they were measured.
search_files, get_file_context and
find_symbol come before any Read. Vanilla gets no tooling
text.
Does adding the code-context MCP server improve the outcome of software-engineering work enough to justify its cost in time and tokens? H0: NV ≤ 0. H1: NV > 0. One-sided, α = 0.05.
| vanilla | cc | |
|---|---|---|
| Model | same (default Opus, recorded per run) | same |
| Prompt | identical task prompt | identical plus the owner's CLAUDE.md code-context rule, verbatim |
| MCP servers | none | code-context only, pre-indexed |
| Hooks, skills, CLAUDE.md, slash commands | none | none |
| Permissions | identical allowlist, sandboxed, offline | identical |
| Workspace | fresh checkout of the pinned commit in an empty temp dir, own git worktree | same |
The added CLAUDE.md rule in the cc prompt is part of the treatment and is reported as such.
claude -p with user, project and local settings sources
disabled, so no CLAUDE.md, user hooks, enabled plugins or personal permissions are
loaded.
| ID | Task | Criterion | Scoring (objective first) |
|---|---|---|---|
| T1 | Feature build in the pinned open-source repository (6 to 10 files across at least 3 modules) | accuracy, time, tokens, code quality | hidden test pass rate; static metrics; blind pairwise judging |
| T2 | 10 systems questions about the codebase ("what breaks if ..."), each needing at least 3 files | IT-systems knowledge | answer key, 0 / 0.5 / 1 per question by a fixed rubric |
| T3 | Stakeholder update about T1 for a non-technical PM | stakeholder communication | blind pairwise rubric: clarity, risk, timeline, no jargon, actionable ask |
| T4 | T1 again, with a requirement change injected after the first plan | control / steerability | extra tokens and time to adapt; hidden tests of the changed requirement |
| S5 | Sequence of 5 related tasks on one codebase, one index | token amortisation | cumulative tokens per task k; break-even k including indexing cost |
Hidden tests and the T2 answer key are never published, so no agent can be trained or prompted against them. Their SHA-256 hashes are published in the headline so that a later disclosure can be checked against what was used.
Q = Σ wi · Δi quality gain C = Σ vj · costj cost increase (negative = cc is cheaper) NV = Q − λ · C λ = 0.5
| Quality criterion | Source | Weight w |
|---|---|---|
| Accuracy | T1 hidden pass rate | 0.30 |
| Code quality | T1: 50 % static score (lint errors, cyclomatic complexity, duplication, coverage of new code), 50 % judge preference | 0.20 |
| Systems knowledge | T2 score | 0.15 |
| Stakeholder communication | T3 judge preference | 0.10 |
| Control | T4: 50 % hidden pass rate after the change, 50 % adaptation cost (tokens plus time, as a cost) | 0.25 |
| Cost criterion | Definition | Weight v |
|---|---|---|
| Wall time | fractional increase of cc over vanilla | 0.5 |
| Total tokens | input + output + cache writes; cache reads at 0.1 weight; fractional increase | 0.5 |
For higher-is-better criteria Δ = (cc − van) / van. For cost criteria the increase is (cc − van) / van.
λ = 0.5 encodes "a quality gain must be at least half the cost increase". The owner's example: a change that is 5 % better but 25 % slower is not worth it.
Q = +0.05 C = +0.25 NV = 0.05 − 0.5 × 0.25 = 0.05 − 0.125 = −0.075 not worth it
The same 5 % gain would be worth it only if the cost increase stayed below 10 %, because then 0.5 × 0.10 = 0.05 and NV is zero.
Every run starts on demand — never from a timer, and nothing is weekly or monthly
scheduled. The legacy Claude runs carry a kind label set at run time
(weekly: 3 repeats per arm per task, T1 to T4; monthly: 10
repeats per arm, plus S5). That label describes how each recorded run was made. It is
not a plan for future runs; each new run records whatever repeat count it was started
with. Results are pushed to the bench-results
branch and this page rebuilds from it. Site code and the method change only through
reviewed pull requests on main.
The code-context workflow is designed around the idea that a human should be able to intercept every step of AI-assisted work, instead of handing a goal to an opaque automated loop and reviewing only the end result.
T4 repeats the T1 feature build, but after the agent has produced its first plan, the requirement changes. Two things are measured: how much of the changed requirement the final code satisfies (hidden tests of the changed requirement), and what the adaptation cost in extra tokens and extra wall time. Each is half of the Control score, which carries weight 0.25 in Q.
T4 measures steerability only in this narrow, automated sense: one injected change, scored objectively. It does not measure the value of human oversight on real projects. The benchmark's headless runs have no human in the loop, so the design benefit described above is a hypothesis about workflow, not something this page claims to have shown. Whether cc is better, worse or no different on T4 is shown only by the Control row in the results table, once data exists.