code-context benchmark

Benchmarks for the code-context MCP server — run on demand, never on a schedule.

LATEST RUN

ALL RUNS

Every published run, newest first, across all kinds. History is append-only: runs are never rewritten or deleted.

RUN DETAIL

METHOD (CLAUDE RUNS)

Pre-registration v2, tasks, judging, statistics, threats to validity
How to read this.
  • The verdict follows the pre-registered rule and uses the 95 % confidence interval of NV, never the point estimate alone.
  • Legacy runs labelled weekly have only n = 3 repeats per arm and task. Their intervals are wide and most of their verdicts are INCONCLUSIVE. That is expected, not a failure.
  • Legacy monthly runs (n = 10, plus S5) are the headline runs; weekly points show the trend only. Runs start on demand — no run is scheduled.
  • NV is a relative figure: +0.10 means a quality gain of 10 percentage points net of the penalised cost. It is not a percentage of money or time.
  • Model and Claude Code versions change between runs. Both are recorded per run; compare trends within one version before reading anything into a change.

This page presents pre-registration v2 (registered 2026-10-09). The authoritative text is PREREGISTRATION.md; v1 is frozen as PREREGISTRATION-v1.md. The version and SHA-256 are recorded in every results row and shown for every run. Changing the method after the first registered run requires a new version with date and reason, and results are reported against the version in force when they were measured.

What changed from v1, and why. v1 was registered and run once (2026-10-09). That run was uninformative, not negative, so v2 changes only what is needed to measure the question. Weights, λ, statistics and the verdict rule are unchanged.
  • Why: the cc arm made 0 code-context calls in 12 of 12 sessions, so the treatment was never received (on a 60-file fixture grep and Read are enough). Both arms also scored at the ceiling (22 of 22 on T1, 11 of 11 on T4).
  • Codebase: a large public open-source repository pinned to one commit (2,000 to 5,000 source files, working offline tests), instead of the 60-file fixture. Repository and commit are recorded in every row.
  • Harder tasks: T1 and T4 touch 6 to 10 files across at least 3 modules; each T2 question needs knowledge spanning at least 3 files. Each task is piloted on the reference solution only, with no arm run.
  • Treatment: the cc prompt carries the owner's real CLAUDE.md rule verbatim: code-context search_files, get_file_context and find_symbol come before any Read. Vanilla gets no tooling text.
  • Treatment-received rule: every session records its code-context calls. A run is VALID only if at least 80 % of cc sessions made at least one. An invalid run is published with its note and excluded from the verdict and the trend, never re-scored or deleted.

Question

Does adding the code-context MCP server improve the outcome of software-engineering work enough to justify its cost in time and tokens? H0: NV ≤ 0. H1: NV > 0. One-sided, α = 0.05.

Arms: only one variable differs

vanilla cc
Model same (default Opus, recorded per run) same
Prompt identical task prompt identical plus the owner's CLAUDE.md code-context rule, verbatim
MCP servers none code-context only, pre-indexed
Hooks, skills, CLAUDE.md, slash commands none none
Permissions identical allowlist, sandboxed, offline identical
Workspace fresh checkout of the pinned commit in an empty temp dir, own git worktree same

The added CLAUDE.md rule in the cc prompt is part of the treatment and is reported as such.

What "vanilla" means exactly

  • Headless claude -p with user, project and local settings sources disabled, so no CLAUDE.md, user hooks, enabled plugins or personal permissions are loaded.
  • Hooks are explicitly emptied. Slash commands and skills are disabled. Session persistence is off.
  • MCP is strict: the vanilla arm has no MCP servers at all, the cc arm has only code-context. Hosted connectors are excluded in both.
  • Both arms use the same permission allowlist (edit inside the worktree, test runners, read-only shell tools) and the same deny list (no web access, no network tools, no pushing, no environment dumps). Shell commands run in an OS sandbox that confines writes to the worktree and blocks network.
  • Both arms start in a fresh empty temp directory with their own git worktree of the same codebase.

Tasks

ID Task Criterion Scoring (objective first)
T1 Feature build in the pinned open-source repository (6 to 10 files across at least 3 modules) accuracy, time, tokens, code quality hidden test pass rate; static metrics; blind pairwise judging
T2 10 systems questions about the codebase ("what breaks if ..."), each needing at least 3 files IT-systems knowledge answer key, 0 / 0.5 / 1 per question by a fixed rubric
T3 Stakeholder update about T1 for a non-technical PM stakeholder communication blind pairwise rubric: clarity, risk, timeline, no jargon, actionable ask
T4 T1 again, with a requirement change injected after the first plan control / steerability extra tokens and time to adapt; hidden tests of the changed requirement
S5 Sequence of 5 related tasks on one codebase, one index token amortisation cumulative tokens per task k; break-even k including indexing cost

Hidden tests and the T2 answer key are never published, so no agent can be trained or prompted against them. Their SHA-256 hashes are published in the headline so that a later disclosure can be checked against what was used.

Judging

  • Objective scores come first: hidden tests, static metrics, the answer key.
  • Where judgement is needed (code quality half, T3), two judges from different non-Claude model families (DeepSeek v4 pro and GLM 5.3) each see both outputs unlabelled, in both orders: 4 verdicts per pair.
  • Preference score = share of verdicts preferring cc, mapped to Δ = 2 × share − 1.
  • A judge that contradicts itself across the two orders counts as a tie for that pair.
  • No Claude model judges, to avoid self-preference.

The formula

Q  = Σ wi · Δi          quality gain
C  = Σ vj · costj       cost increase (negative = cc is cheaper)
NV = Q − λ · C         λ = 0.5
Quality criterion Source Weight w
Accuracy T1 hidden pass rate 0.30
Code quality T1: 50 % static score (lint errors, cyclomatic complexity, duplication, coverage of new code), 50 % judge preference 0.20
Systems knowledge T2 score 0.15
Stakeholder communication T3 judge preference 0.10
Control T4: 50 % hidden pass rate after the change, 50 % adaptation cost (tokens plus time, as a cost) 0.25
Cost criterion Definition Weight v
Wall time fractional increase of cc over vanilla 0.5
Total tokens input + output + cache writes; cache reads at 0.1 weight; fractional increase 0.5

For higher-is-better criteria Δ = (cc − van) / van. For cost criteria the increase is (cc − van) / van.

The owner's rule, worked

λ = 0.5 encodes "a quality gain must be at least half the cost increase". The owner's example: a change that is 5 % better but 25 % slower is not worth it.

Q  = +0.05
C  = +0.25
NV = 0.05 − 0.5 × 0.25 = 0.05 − 0.125 = −0.075   not worth it

The same 5 % gain would be worth it only if the cost increase stayed below 10 %, because then 0.5 × 0.10 = 0.05 and NV is zero.

Statistics

  • Paired design: each repeat runs both arms on the same task; arm order is randomised.
  • For every Δ and for NV: the mean and a 95 % bias-corrected bootstrap CI (10,000 resamples over pairs).
  • Wilcoxon signed-rank test per criterion, effect size r, Holm correction across criteria.
  • Every run is reported, including failed and timed-out runs (scored 0, cost as measured up to the timeout). No run is excluded after the fact, except that a run whose treatment-received rate is below 80 % is marked INVALID, published with its note and left out of the verdict and the trend.

Verdict rule

  • KEEP if the lower bound of NV's 95 % CI is above 0.
  • DROP if the upper bound is below 0.
  • INCONCLUSIVE otherwise.
  • Never a verdict from a point estimate alone.

Cadence

Every run starts on demand — never from a timer, and nothing is weekly or monthly scheduled. The legacy Claude runs carry a kind label set at run time (weekly: 3 repeats per arm per task, T1 to T4; monthly: 10 repeats per arm, plus S5). That label describes how each recorded run was made. It is not a plan for future runs; each new run records whatever repeat count it was started with. Results are pushed to the bench-results branch and this page rebuilds from it. Site code and the method change only through reviewed pull requests on main.

Control and steerability

The code-context workflow is designed around the idea that a human should be able to intercept every step of AI-assisted work, instead of handing a goal to an opaque automated loop and reviewing only the end result.

  • The plan is visible. Work is broken into tickets, sprints and phases that exist outside the model's context, so you can read what is about to happen.
  • Gates sit between phases. A phase does not start until the previous one is accepted, which gives natural places to stop, correct or redirect.
  • Decisions are logged. Choices and their reasons are written down, so a change of direction does not depend on what the model still remembers.
  • An opaque loop optimises for finishing. It can be faster and cheaper when the first plan is right. The open question is what happens when the plan is wrong.

What T4 measures

T4 repeats the T1 feature build, but after the agent has produced its first plan, the requirement changes. Two things are measured: how much of the changed requirement the final code satisfies (hidden tests of the changed requirement), and what the adaptation cost in extra tokens and extra wall time. Each is half of the Control score, which carries weight 0.25 in Q.

T4 measures steerability only in this narrow, automated sense: one injected change, scored objectively. It does not measure the value of human oversight on real projects. The benchmark's headless runs have no human in the loop, so the design benefit described above is a hypothesis about workflow, not something this page claims to have shown. Whether cc is better, worse or no different on T4 is shown only by the Control row in the results table, once data exists.

Threats to validity

  • Single codebase (external validity). One pinned open-source repository. Results may not transfer to other languages, sizes or domains.
  • LLM judges. Mitigated by two model families, both presentation orders, no Claude judge, and objective scores first. Judges can still share biases.
  • Model and Claude Code version drift. Both are recorded per row. A change in either can move results independently of code-context; trends are shown per version.
  • The cc prompt rule. The cc arm's prompt carries the owner's CLAUDE.md rule. That difference is part of the treatment and cannot be separated from the server itself.
  • Small n. Legacy weekly runs have n = 3, which gives wide CIs; the monthly n = 10 runs are the headline runs. Runs start on demand, so n is whatever the started run recorded.