What Does a Harness Buy? Tokens, Mostly.

Summary (Overview)

  • Core finding: The coding agent harness (system prompt, tools, context management) barely moves pass rate on SWE-bench Verified when the model is held fixed—swapping harnesses flips as many tasks as simply rerunning the same harness (13% at the median).
  • The one measurable harness effect is a loss, not a gain: OpenCode trails by up to 9 points on the large pool, with half that gap attributable to runs cut short by its output cap with no recovery mechanism.
  • The harness decisively sets the bill: Cost per task differs by up to 3× across harnesses with the same model, driven primarily by the fixed preamble (system prompt + tool schemas) sent with every step—16k tokens in Claude Code, 7k in OpenCode, under 1k in mini-SWE-agent.
  • Statistical resolution limits: 45 tasks catch a 13-point gap only half the time and no gap at 80% power; 447 tasks resolve about 5 points—coarser than many published harness gains.
  • No vendor home-field advantage: DeepSeek's own harness on its own model scored inside the range of the other three harnesses.

Introduction and Theoretical Foundation

Coding agents have become a product category defined as model + harness. The harness (also called a scaffold) comprises everything around the model: the system prompt, available tools, feedback of tool output, and conversation trimming. Production harnesses ship more than two releases daily, and the field widely believes harness engineering moves scores.

Published evidence for harness effects includes:

  • Same model scoring 52.1% under one harness vs. 57.8% under another (Vats & Golev, 2026)
  • Six harnesses over the same model pool spanning 23.8 points (Yao et al., 2026)
  • Five harnesses on one fixed model spanning 27.4 points (Zheng et al., 2026)

Two problems make these numbers hard to read:

  1. Almost all come from one run per cell, while ten reruns of one configuration on SWE-bench Verified spread by 2.2 to 6.0 points (Bjarnason et al., 2026)
  2. Harnesses are often not on equal footing—one study shows a harness scoring 19.1% with a minimal benchmark adapter vs. 73.4% with the full adapter on the same backbone

The central question: Does swapping the harness move a score more than running the same harness again? This had not been measured on the same tasks and models until this work.

Methodology

Harnesses (versions pinned)

HarnessVersionDesign
Claude Code2.1 (releases 2.1.259–2.1.273)Typed file/search/edit tools, long system prompt, automatic compaction
mini-SWE-agent2.4.6Single bash tool, no context management, hard stop at window
OpenCode1.18 (1.18.30–1.18.31)Typed tools, long system prompt, automatic compaction

Models

  • Local (vLLM 0.27.1, bfloat16): Qwen3.6-35B-A3B, Qwen3.8-27B
  • API: DeepSeek-V4-Flash (release 0731), GLM-5.3-Flash, HY4-Preview
  • Anchor: Claude Opus 5 (inside Claude Code only)

Task Sets (from SWE-bench Verified, 492 usable tasks)

  • hard45: All 45 tasks in the two hardest difficulty bands (1–4 hours and 4+ hours)
  • pool447: The remaining 447 tasks

Design

  • One step = one model call; every trial gets 300 steps
  • All runs offline (network isolation enforced after an audit found 71% of web-enabled runs fetched upstream fixes)
  • Task-level outcomes paired across harnesses; gaps tested with McNemar's exact test, Newcombe intervals
  • Equivalence at δ = 5 points tested with two one-sided tests (TOST)
  • Reruns of identical configurations used only to measure rerun variance
  • Cost computed from each trial's own token counts under vendor list prices

Empirical Validation / Results

4.1 Pass Rate: The Harness Barely Moves It

pool447 results (Table 1): Claude Code vs. mini-SWE-agent (heaviest vs. lightest harness) are equivalent within ±5 points on both models:

ModelPairΔ (pp)95% CIVerdict
Qwen3.6-35B-A3BCC − mini−1.8[−5.2, +1.6]equivalent
Qwen3.6-35B-A3BCC − OC+2.9[−0.9, +6.7]inconclusive
Qwen3.6-35B-A3Bmini − OC+4.7[+1.1, +8.4]inferior B
Qwen3.8-27BCC − mini−1.4[−4.2, +1.5]equivalent
Qwen3.8-27BCC − OC+7.9[+4.3, +11.5]inferior B
Qwen3.8-27Bmini − OC+9.2[+5.5, +12.9]inferior B

hard45 results (Table 2): Harness spread is 2–5 tasks per model, comparable to rerun spread (up to 3 tasks). Mini-SWE-agent leads on 3 of 5 models, Claude Code on the other 2. All paired differences have intervals covering zero.

ModelClaude Codemini-SWE-agentOpenCode
Qwen3.6-35B-A3B11/45 (24%)15/45 (33%)11/45 (24%)
Qwen3.8-27B19/43 (44%)17/45 (38%)17/45 (38%)
DeepSeek-V4-Flash21/45 (47%)22/45 (49%)19/45 (42%)
GLM-5.3-Flash22/45 (49%)20/45 (44%)18/45 (40%)
HY4-Preview23/45 (51%)26/45 (58%)21/45 (47%)
Claude Opus 536/45 (80%)——

4.2 The Noise Floor

  • Flip rate (share of tasks whose outcome differs): reruns flip 13% at the median; harness swaps flip the same 13%; model swaps flip 22%
  • Task-level flips don't replicate: On harness pairs run twice, 5 tasks split the same way and 7 the opposite way (coin flip). Model swaps under the same harness show 10 of 15 tasks splitting the same way both times
  • Resolution limits: With 14% discordance, 45 tasks catch a 12.9-point gap half the time, nothing at 80% power; 447 tasks detect 5.2 points at 80% power

4.3 Harnesses Lose Tasks

OpenCode's output cap failure: When model reasoning exceeds the provider's cap mid-step, OpenCode doesn't re-prompt. On pool447 with Qwen3.8-27B, 33 OpenCode trials fail this way vs. 1 for Claude Code (which re-prompts). Removing cutoff-affected tasks roughly halves OpenCode's gap.

OpenCode stops sooner: On Qwen3.6-35B-A3B, OpenCode ends after median 28 model calls vs. 71 for mini-SWE-agent and 55 for Claude Code.

Other failure modes:

  • mini-SWE-agent overflows context window (no compaction): 11/45 hard tasks on Qwen3.8-27B, but most are tasks the model was already losing
  • OpenCode hits wall-clock limits (had 1/6 of Claude Code's time on local arms)
  • OpenCode's search tools can't run offline

4.4 The Harness Sets the Bill

Cost per task differs by up to 3× across harnesses on the same model. The bill decomposes into:

Bill=per-step input×steps×price per tokenBill = \text{per-step input} \times \text{steps} \times \text{price per token}

Preamble sizes (system prompt + tool schemas, paid every step):

  • Claude Code: 16,581 tokens
  • OpenCode: 7,025 tokens
  • mini-SWE-agent: 829 tokens

Per-step input growth (Figure 4): The harnesses separate at the first call—on GLM-5.3-Flash, first call carries 17.1k tokens (Claude Code), 7.6k (OpenCode), 1.4k (mini-SWE-agent). Slopes sit within 1.5× of each other.

Cache pricing scales but doesn't reorder: 96–99% of input tokens are cache hits. Zhipu charges 0.287× miss price for hits (cached input dominates GLM bill); DeepSeek charges 0.033× (cached transcript is under half the bill). Repricing under either vendor's list keeps the harness ordering.

4.5 Non-Findings

  • No home-field advantage: DeepSeek's harness on DeepSeek-V4-Flash passes 21/45, same as Claude Code, at 1.04 CNY vs. 0.98 CNY per trial
  • No score gain from tool ablations or old releases: Removing Claude Code's file-edit tools or shell changes nothing; a year-old release (1.0.100) passes within one task of control at 0.44× cost
  • Thinking interaction: Turning off thinking on Qwen3.6-35B-A3B costs mini-SWE-agent 11 points (p = .04) but moves Claude Code and OpenCode by less than noise

Theoretical and Practical Implications

  1. Harness engineering is loss-avoidance, not win-creation: Every measurable harness effect is a way to lose a task (output caps without recovery, context overflows without compaction, tools that fail offline). No harness was found to win tasks another harness loses.

  2. The dominant harness effect is economic: The 3× cost spread is set at the first call by the preamble and multiplied by step count. The most expensive harness (Claude Code) buys no measurable pass-rate advantage over the cheapest (mini-SWE-agent).

  3. Statistical practice matters: Most published harness comparisons run one cell once and report gains below the resolution their sample size allows. The cheapest control is a second run of the same cell, which most comparisons skip.

  4. Cache pricing interacts with harness design: Providers with cheap cache hits (DeepSeek at 0.033×) make the verbose harness competitive on cost; providers with expensive hits (Zhipu at 0.287×) amplify the verbose harness's bill.

Conclusion

Main takeaways:

  • On SWE-bench Verified, swapping harnesses under a fixed model moves pass rate by as much as rerunning the same harness on the hardest 45 tasks
  • The only harness effects that clear the noise are losses (OpenCode trailing by 8–9 points on Qwen3.8-27B, tied to trial-ending behavior)
  • The harness decisively sets the bill—up to 3× on the same model—through preamble size, per-step growth, and step count
  • The key numbers to carry between vendors: first-call size, growth per step, and step count

Future directions:

  • The thinking interaction is the lead worth chasing: thinking dependence appears specific to model–harness pairs, not models alone
  • A recovery step for OpenCode's output cutoff (which Claude Code already has) is the one change that might move a score
  • Need transfer checks beyond SWE-bench (Terminal-Bench 2.1 checked as transfer); general-purpose agents may be a different regime
  • The closed anchor model (Claude Opus 5) ran only in Claude Code, so the three-harness ladder couldn't test whether frontier models show the same gaps

Recommendation: Report a second run of the same cell next to any harness gap—it is the cheapest control and the one most comparisons skip.

Related papers