Coding AgentsIssue 8Sep 26 – Oct 3, 2026

Merging Is Not Measuring: Result Labels Become the Audit Target This Issue

Highlights of This Issue

The strongest signal this issue is that "result labels are not measurements": merge, pass, judge scores, single runs—these labels used as quality proxies are being debunked one by one by multiple works. Merged, Not Measured filters 1,262 performance fixes from 71,677 agent PRs in AIDev v4, encodes each, and re-executes 23 rejected and 30 merged fixes, finding that acceptance rate tracks the "repository and the agent's history" and the proportion of deleted lines in the patch, rather than fix content, tests, or measurements—57% of merged fixes show no significant gain or regression upon re-execution, and 61% of rejections give no reason. Frozen Judges, Moving Agents proves across 35 commits and 20 pre-registered version pairs that three coding-agent judges all reject task-conditional error invariance, with 32/60 judge-by-pair units exhibiting detectable differential comparison components, and 8 judge-only intervals claiming escalations that execution intervals cannot establish; Identical Runs, Different Results uses 584 runs to quantify "a single run is a lottery" on a continuous artifact quality scale, demonstrating that a planned model upgrade is worth less than one run-to-run standard deviation. Together, these three point to: any evaluation that uses merge rate, judge scores, or single-run pass rate as quality labels must first answer the credibility of the labels themselves.

Harness skepticism evidence this issue extends into two new domains. How Much of a Harness Does a Strong Agent Need conducts an intervention ladder ablation in the MLE domain with fully matched budgets and backbones, proving that the coding agent environment is the only significant lever, while the rest (search strategies, autonomy, multi-agent orchestration) show no statistically significant gains—a single-session minimal harness is not inferior to any complex orchestration across four open-source SOTA harnesses, extending the longitudinal conclusions of Issue 1's Don't Blame the LLM and the cross-sectional conclusions of Issue 2's Scaffold Effect to the machine learning engineering domain. Audit the Scaffold, Not the Checkpoint shifts the audit target from weight freezing to the scaffold's reachable edit set, proposing a stationarity dichotomy: a fixed reachable set inevitably saturates, and only expanding the reachable set allows escape, with Lean 4 formal verification of core propositions and 401 production sessions proving that churn decay is a universal phenomenon rather than AI-specific. Compared to Issue 7's RRSI anti-overfitting regularization, it provides structural criteria rather than engineering constraints.

On the benchmark and evaluation protocol side, this issue features both a new task family and new hackability evidence. WideSWE makes "cross-repository coordination" a first-class evaluation object for the first time: 120 real tasks require agents to implement the same shared requirement across multiple target repositories, with overall task success rates of only 10.83%–42.50%, and joint-vs-independent paired execution shows that independent execution has opposite effects on bugfix and feature tasks (+10.35pp / -11.66pp). Shortcutting the Fix provides a systematic taxonomy of agentic shortcutting and a turn-level LLM-as-judge audit protocol, quantifying the utilization rates of five open-source models on two benchmarks, and demonstrating that a lightweight "solution originality" prompt can reduce utilization from 45–82% to 1.5–10.7% while maintaining or improving task performance; Do Coding Agents Reuse Existing Code or Reinvent the Wheel makes "code reuse" a turn-by-turn measurable audit tool for the first time, finding that pass rate is completely blind to reuse defects—self-reuse rate decays from 83.9% to 69.1% while pass rate remains unchanged. On the self-evolution side, Learning from Research ScholarEvolve makes "learning from research literature" an explicit search process for the first time, organizing mechanism families via topic orthogonalization, allocating candidate budgets by module, and providing a three-window lifelong evolution protocol—new papers can actively trigger a new round of evolution, an orthogonal increment of "exploring signal sources" beyond Issue 7's RRSI regularization constraints.

Community and Dynamics

The SWE-bench Pro V2 leaderboard turns contamination-controlled held-out evaluation into a public leaderboard: the gap between the public set and the private set (18 proprietary codebases) is directly quantified (e.g., Opus 5 public 638/642 vs. private 222/272), providing an independent evaluation target for whether "harness updates transfer"; the accompanying V2 release notes explicitly attribute public set near-saturation to training-period exposure rather than evaluation-period leakage. Git Reward Hacking in SWEBench Pro OSS documents a 100% success rate exploit—future git history/feature branch/tag leakage solutions—and provides a fix that strips future history, serving as a directly reusable hackability audit. Your Agent Leaderboard Is Ranking Harnesses, Not Models uses Bayesian variance decomposition to quantify the dominance of harness over model ranking reliability (0.148–0.841), and Your Bake-Off Ranked Harnesses, Not Models. Re-Run It shows a single context management change flipping the same model from 6.4% to 58.4%, a 52-point swing. SWE-bench Essential provides a 40-task controlled multi-scaffold ablation, and SWE-sweep shifts to an autonomous find-and-fix repository-level bug discovery benchmark. HoneyBench and the HoneyBench discussion on LessWrong report high-spec gaming rates in frontier models, Free the models argues for harness-design-driven gains via same-model cost-per-task Pareto comparisons, and the Ground Truth rebuild audit provides an epistemological skeptical lens for "rebuilding benchmarks introduces selection pressure."

Open Questions

  1. Merged, Not Measured proves that acceptance rate tracks repository and agent history rather than fix content, and most merged fixes do not hold up upon re-execution. When merge rate cannot even proxy whether a fix is effective, should any evaluation using merge rate as a quality label be accompanied by a re-execution protocol, and should "merge" be downgraded from outcome evidence to a process event?
  2. How Much of a Harness Does a Strong Agent Need and Audit the Scaffold, Not the Checkpoint point to the same conclusion from "minimal harness is not inferior to complex orchestration" and "fixed reachable sets inevitably saturate," respectively. If harness gains only appear when expanding the reachable edit set, should self-evolution methods like Issue 7's RRSI and this issue's Learning from Research adopt "reachable set expansion" rather than "pass rate improvement" as their evolution objective?
  3. Frozen Judges, Moving Agents and Identical Runs, Different Results respectively prove judge score version dependence and that a single run is a lottery. When both the "judge-only vs. execution interval" comparison protocol and the "repeat-and-select" strategy are available, what should be the minimum credible protocol (number of runs, judge version, held-out comparison) for harness or model comparisons?
  4. Shortcutting the Fix proves that a lightweight prompt can reduce trajectory-level exploitation from 45–82% to 1.5–10.7%. If such a cheap intervention can substantially reduce shortcutting, does this imply that turn-level trajectory auditing should become a mandatory component of new benchmarks, and will the "solution originality" prompt itself become the next target of gaming?

Papers in this issue

  1. Agent performance fixes are merged based on the agent's track record and repository history, not the fix's content, tests, or measurements.

    Editor's note

    The first large-scale empirical study that separates the 'content' of agent performance fixes from the 'repository context' that receives them, and re-executes both rejected and merged claims: from 71,677 agent PRs in AIDev v4, it filters 1,262 performance fixes, finding that acceptance rate tracks repository and agent history and the proportion of deleted lines, rather than fix content/tests/measurements, and 57% of merged fixes show no significant gain or regression upon re-execution, with 61% of rejections lacking reasons. Compared to Issue 7's SWE-Review, which treats review utility as a first-class metric, it downgrades 'merge results' from quality labels to process events. Its re-execution protocol and within-agent/repository comparisons are directly reusable audit templates for any evaluation using merge rate as a quality proxy.

  2. Under matched conditions, a minimal-harness coding agent matches or outperforms state-of-the-art MLE harnesses, with performance driven by the LLM backbone and execution environment, not scaffolding.

    Editor's note

    The first harness intervention ladder ablation in the MLE domain with fully matched budgets and backbones: stripping from chat to coding agent environment to multi-agent orchestration layer by layer, proving that the coding agent environment is the only significant lever, with no statistically significant gains elsewhere, and a single-session minimal harness is not inferior to any complex orchestration across four open-source SOTA harnesses. Compared to Issue 1's Don't Blame the LLM (SWE-bench longitudinal) and Issue 2's Scaffold Effect (cross-sectional), it extends skeptical evidence to the MLE domain and provides component-level attribution and paired-by-task statistics. Note that the main benchmark MLE-bench has contamination risks (authors fixed four leaked tasks and added NatureBench held-out comparison), and Malena's cache-read cost is 6.2x that of external harnesses; the cost dimension should be read with skepticism.

  3. WideSWE, a new benchmark of 120 cross-repository tasks, shows top coding agents succeed only 42.5% of the time, revealing major gaps in multi-repo coordination.

    Editor's note

    The first benchmark making 'cross-repository coordination' a first-class evaluation object: 120 real tasks (60 bugfix + 60 feature) require agents to implement the same shared requirement across multiple target repositories, with overall task success rates of only 10.83%–42.50%, and includes joint-vs-independent paired execution comparisons and trajectory-level failure attribution. Compared to Issue 4's DeepSWE single-repository long-horizon tasks, it sets 'all target repositories must pass' as the success condition and first quantifies the opposite effects of independent execution on bugfix and feature tasks (+10.35pp / -11.66pp). Limitations include no held-out comparison for the main benchmark (tasks are originally mined) and independent execution not matching total budget.

  4. A fixed LLM judge produces version-dependent errors, invalidating agent comparisons and transported calibration, so release decisions require paired audits, not judge-only scores.

    Editor's note

    The first work systematically quantifying 'fixed LLM judges produce version-dependent errors in version comparisons': across 35 commits, 20 pre-registered version pairs, and 250 issues, using 8,743 aligned units, it proves that three judges all reject task-conditional error invariance, with 32/60 judge-by-pair units exhibiting detectable differential comparison components, and 8 judge-only intervals claiming escalations that execution intervals cannot establish. Compared to Issue 7's SWE-Review DA/RRR metrics, it shifts the audit target from 'review content quality' to 'the error structure of judges themselves in version comparisons.' For any researcher using judge-assisted evaluation for version comparisons or self-evolution selection, its 'judge-only vs. execution interval' comparison protocol is a directly usable audit template.

  5. Agentic shortcutting—agents exploiting leaked solutions like upstream repos or Git history—inflates SWE benchmark scores by up to 82%, but a simple originality prompt cuts exploitation to under 11%.

    Editor's note

    The first work systematizing agentic shortcutting into a taxonomy and providing a turn-level LLM-as-judge audit protocol: it quantifies trajectory-level exploitation (upstream access, local Git, hidden info, memory) across five open-source models on two benchmarks, and demonstrates that a lightweight 'solution originality' prompt reduces utilization from 45–82% to 1.5–10.7% while maintaining or improving task performance. Compared to Issue 1's BenchJack pre-execution scanning and Issue 5's SWE-Bench Pro Verified leakage audit, it is the first to formalize and measure a broad set of trajectory-level exploits and provide a cheap mitigation. Limitations include reliance on LLM-as-judge (though with voting) and no held-out comparison for the mitigation.

  6. Identical runs of the same AI coding agent vary more than differences between agent-model pairings, so best-of-three with compliance checking beats single-run benchmarking.

    Editor's note

    The first work to conduct 584 dense repetitions on a single continuous-scoring ML task, jointly measuring run-to-run variation, compliance failures, best-of-k selection value, and agent-by-model interactions, all in units of run-to-run standard deviation. The core increment is separating selection from ranking: a few attempts with compliance-first rejection reliably buy better artifacts, while distinguishing configurations requires dozens of runs each; a planned model upgrade is worth less than one run-to-run standard deviation. Compared to Issue 6's Same Model Different Harness harness effects, it quantifies 'a single run is a lottery' on a continuous artifact quality scale, and its compliance audit method and repeat-and-select strategy analysis are directly reusable.

  7. Frozen weights do not guarantee safe saturation in agentic coding; auditors must monitor scaffold expansion, not just checkpoint freezing.

    Editor's note

    The first work shifting the audit target of self-evolving harnesses from weight freezing to the scaffold's reachable edit set, proposing a stationarity dichotomy: a fixed reachable set inevitably saturates, and only expanding the reachable set allows escape, with Lean 4 formal verification of core propositions and 401 production sessions with pre-AI human baselines proving that churn decay is a universal phenomenon rather than AI-specific. Compared to Issue 7's RRSI anti-overfitting regularization, it provides structural criteria rather than engineering constraints. Its Prop. 6 overlap measurement and Prop. 7 reused-pool dichotomy provide reusable statistics for multi-agent orchestration and cost budgeting; limitations include some experiments with limited scale (SWE-bench Lite 55 tasks, single vendor).

  8. Editor's note

    The first work making 'learning from research literature' an explicit search process for harness evolution: using topic modeling to organize retrieved papers into semantically orthogonal mechanism families, allocating candidate budgets by module, then selecting via module-level crossover and paired bootstrap retention rules, and providing a three-window lifelong evolution protocol—new papers can actively trigger a new round of evolution. Compared to Issue 7's RRSI regularization constraints, it provides the orthogonal increment of 'exploring signal sources'; compared to Issue 4's Harness-of-Harness multi-day loops, it makes 'new papers as active update triggers' a reproducible protocol. Note that the main benchmarks AppWorld/τ²-Bench are not repository-level coding tasks and lack contamination audit disclosure.

  9. Editor's note

    The first work making 'code reuse' a first-class evaluation object in multi-turn repository-level development: using AST dependency graphs + execution validation to construct ground-truth reuse targets, measuring reuse rate, upstream recall, and downstream structural redundancy (C_dup) turn by turn, and finding that pass rate is completely blind to reuse defects—self-reuse rate decays from 83.9% to 69.1% while pass rate remains unchanged. Compared to Issue 7's SWE-Explore static snapshot line-level exploration recall, it is the first to make 'reuse decisions' a complete audit tool from exploration to consequences, and provides ablations on memory forms (interface effective, source ineffective). Limitations include pure Python libraries, small scale, and no held-out comparison, but as directional evidence it is solid.