What Does a Harness Buy? Tokens, Mostly.
Summary (Overview)
- Core finding: The coding agent harness (system prompt, tools, context management) barely moves pass rate on SWE-bench Verified when the model is held fixed—swapping harnesses flips as many tasks as simply rerunning the same harness (13% at the median).
- The one measurable harness effect is a loss, not a gain: OpenCode trails by up to 9 points on the large pool, with half that gap attributable to runs cut short by its output cap with no recovery mechanism.
- The harness decisively sets the bill: Cost per task differs by up to 3× across harnesses with the same model, driven primarily by the fixed preamble (system prompt + tool schemas) sent with every step—16k tokens in Claude Code, 7k in OpenCode, under 1k in mini-SWE-agent.
- Statistical resolution limits: 45 tasks catch a 13-point gap only half the time and no gap at 80% power; 447 tasks resolve about 5 points—coarser than many published harness gains.
- No vendor home-field advantage: DeepSeek's own harness on its own model scored inside the range of the other three harnesses.
Introduction and Theoretical Foundation
Coding agents have become a product category defined as model + harness. The harness (also called a scaffold) comprises everything around the model: the system prompt, available tools, feedback of tool output, and conversation trimming. Production harnesses ship more than two releases daily, and the field widely believes harness engineering moves scores.
Published evidence for harness effects includes:
- Same model scoring 52.1% under one harness vs. 57.8% under another (Vats & Golev, 2026)
- Six harnesses over the same model pool spanning 23.8 points (Yao et al., 2026)
- Five harnesses on one fixed model spanning 27.4 points (Zheng et al., 2026)
Two problems make these numbers hard to read:
- Almost all come from one run per cell, while ten reruns of one configuration on SWE-bench Verified spread by 2.2 to 6.0 points (Bjarnason et al., 2026)
- Harnesses are often not on equal footing—one study shows a harness scoring 19.1% with a minimal benchmark adapter vs. 73.4% with the full adapter on the same backbone
The central question: Does swapping the harness move a score more than running the same harness again? This had not been measured on the same tasks and models until this work.
Methodology
Harnesses (versions pinned)
| Harness | Version | Design |
|---|---|---|
| Claude Code | 2.1 (releases 2.1.259–2.1.273) | Typed file/search/edit tools, long system prompt, automatic compaction |
| mini-SWE-agent | 2.4.6 | Single bash tool, no context management, hard stop at window |
| OpenCode | 1.18 (1.18.30–1.18.31) | Typed tools, long system prompt, automatic compaction |
Models
- Local (vLLM 0.27.1, bfloat16): Qwen3.6-35B-A3B, Qwen3.8-27B
- API: DeepSeek-V4-Flash (release 0731), GLM-5.3-Flash, HY4-Preview
- Anchor: Claude Opus 5 (inside Claude Code only)
Task Sets (from SWE-bench Verified, 492 usable tasks)
- hard45: All 45 tasks in the two hardest difficulty bands (1–4 hours and 4+ hours)
- pool447: The remaining 447 tasks
Design
- One step = one model call; every trial gets 300 steps
- All runs offline (network isolation enforced after an audit found 71% of web-enabled runs fetched upstream fixes)
- Task-level outcomes paired across harnesses; gaps tested with McNemar's exact test, Newcombe intervals
- Equivalence at δ = 5 points tested with two one-sided tests (TOST)
- Reruns of identical configurations used only to measure rerun variance
- Cost computed from each trial's own token counts under vendor list prices
Empirical Validation / Results
4.1 Pass Rate: The Harness Barely Moves It
pool447 results (Table 1): Claude Code vs. mini-SWE-agent (heaviest vs. lightest harness) are equivalent within ±5 points on both models:
| Model | Pair | Δ (pp) | 95% CI | Verdict |
|---|---|---|---|---|
| Qwen3.6-35B-A3B | CC − mini | −1.8 | [−5.2, +1.6] | equivalent |
| Qwen3.6-35B-A3B | CC − OC | +2.9 | [−0.9, +6.7] | inconclusive |
| Qwen3.6-35B-A3B | mini − OC | +4.7 | [+1.1, +8.4] | inferior B |
| Qwen3.8-27B | CC − mini | −1.4 | [−4.2, +1.5] | equivalent |
| Qwen3.8-27B | CC − OC | +7.9 | [+4.3, +11.5] | inferior B |
| Qwen3.8-27B | mini − OC | +9.2 | [+5.5, +12.9] | inferior B |
hard45 results (Table 2): Harness spread is 2–5 tasks per model, comparable to rerun spread (up to 3 tasks). Mini-SWE-agent leads on 3 of 5 models, Claude Code on the other 2. All paired differences have intervals covering zero.
| Model | Claude Code | mini-SWE-agent | OpenCode |
|---|---|---|---|
| Qwen3.6-35B-A3B | 11/45 (24%) | 15/45 (33%) | 11/45 (24%) |
| Qwen3.8-27B | 19/43 (44%) | 17/45 (38%) | 17/45 (38%) |
| DeepSeek-V4-Flash | 21/45 (47%) | 22/45 (49%) | 19/45 (42%) |
| GLM-5.3-Flash | 22/45 (49%) | 20/45 (44%) | 18/45 (40%) |
| HY4-Preview | 23/45 (51%) | 26/45 (58%) | 21/45 (47%) |
| Claude Opus 5 | 36/45 (80%) | — | — |
4.2 The Noise Floor
- Flip rate (share of tasks whose outcome differs): reruns flip 13% at the median; harness swaps flip the same 13%; model swaps flip 22%
- Task-level flips don't replicate: On harness pairs run twice, 5 tasks split the same way and 7 the opposite way (coin flip). Model swaps under the same harness show 10 of 15 tasks splitting the same way both times
- Resolution limits: With 14% discordance, 45 tasks catch a 12.9-point gap half the time, nothing at 80% power; 447 tasks detect 5.2 points at 80% power
4.3 Harnesses Lose Tasks
OpenCode's output cap failure: When model reasoning exceeds the provider's cap mid-step, OpenCode doesn't re-prompt. On pool447 with Qwen3.8-27B, 33 OpenCode trials fail this way vs. 1 for Claude Code (which re-prompts). Removing cutoff-affected tasks roughly halves OpenCode's gap.
OpenCode stops sooner: On Qwen3.6-35B-A3B, OpenCode ends after median 28 model calls vs. 71 for mini-SWE-agent and 55 for Claude Code.
Other failure modes:
- mini-SWE-agent overflows context window (no compaction): 11/45 hard tasks on Qwen3.8-27B, but most are tasks the model was already losing
- OpenCode hits wall-clock limits (had 1/6 of Claude Code's time on local arms)
- OpenCode's search tools can't run offline
4.4 The Harness Sets the Bill
Cost per task differs by up to 3× across harnesses on the same model. The bill decomposes into:
Preamble sizes (system prompt + tool schemas, paid every step):
- Claude Code: 16,581 tokens
- OpenCode: 7,025 tokens
- mini-SWE-agent: 829 tokens
Per-step input growth (Figure 4): The harnesses separate at the first call—on GLM-5.3-Flash, first call carries 17.1k tokens (Claude Code), 7.6k (OpenCode), 1.4k (mini-SWE-agent). Slopes sit within 1.5× of each other.
Cache pricing scales but doesn't reorder: 96–99% of input tokens are cache hits. Zhipu charges 0.287× miss price for hits (cached input dominates GLM bill); DeepSeek charges 0.033× (cached transcript is under half the bill). Repricing under either vendor's list keeps the harness ordering.
4.5 Non-Findings
- No home-field advantage: DeepSeek's harness on DeepSeek-V4-Flash passes 21/45, same as Claude Code, at 1.04 CNY vs. 0.98 CNY per trial
- No score gain from tool ablations or old releases: Removing Claude Code's file-edit tools or shell changes nothing; a year-old release (1.0.100) passes within one task of control at 0.44× cost
- Thinking interaction: Turning off thinking on Qwen3.6-35B-A3B costs mini-SWE-agent 11 points (p = .04) but moves Claude Code and OpenCode by less than noise
Theoretical and Practical Implications
-
Harness engineering is loss-avoidance, not win-creation: Every measurable harness effect is a way to lose a task (output caps without recovery, context overflows without compaction, tools that fail offline). No harness was found to win tasks another harness loses.
-
The dominant harness effect is economic: The 3× cost spread is set at the first call by the preamble and multiplied by step count. The most expensive harness (Claude Code) buys no measurable pass-rate advantage over the cheapest (mini-SWE-agent).
-
Statistical practice matters: Most published harness comparisons run one cell once and report gains below the resolution their sample size allows. The cheapest control is a second run of the same cell, which most comparisons skip.
-
Cache pricing interacts with harness design: Providers with cheap cache hits (DeepSeek at 0.033×) make the verbose harness competitive on cost; providers with expensive hits (Zhipu at 0.287×) amplify the verbose harness's bill.
Conclusion
Main takeaways:
- On SWE-bench Verified, swapping harnesses under a fixed model moves pass rate by as much as rerunning the same harness on the hardest 45 tasks
- The only harness effects that clear the noise are losses (OpenCode trailing by 8–9 points on Qwen3.8-27B, tied to trial-ending behavior)
- The harness decisively sets the bill—up to 3× on the same model—through preamble size, per-step growth, and step count
- The key numbers to carry between vendors: first-call size, growth per step, and step count
Future directions:
- The thinking interaction is the lead worth chasing: thinking dependence appears specific to model–harness pairs, not models alone
- A recovery step for OpenCode's output cutoff (which Claude Code already has) is the one change that might move a score
- Need transfer checks beyond SWE-bench (Terminal-Bench 2.1 checked as transfer); general-purpose agents may be a different regime
- The closed anchor model (Claude Opus 5) ran only in Claude Code, so the three-harness ladder couldn't test whether frontier models show the same gaps
Recommendation: Report a second run of the same cell next to any harness gap—it is the cheapest control and the one most comparisons skip.
Related papers
- Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild
Agent harnesses remain substantially undertested, with less than half of LLM-dependent code covered, and HarnessTester's contract-aware test generation boosts coverage and mutation scores by up to 95% while finding 88 previously-unknown bugs.
- Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs
Hybrid language models over-rely on attention and underuse recurrent memory, but an auxiliary loss forcing recurrent routing improves QA accuracy by up to 5.2% and agentic success rates.
- VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
VERSE shows LLM optimizers improve by evolving their own harness, but only when execution-based verification tools are provided, boosting SWE-rebench accuracy across all baselines.