Full text not available for this paper
Summary of "Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding"
Summary (Overview)
-
Core theoretical contribution: The paper establishes a stationarity dichotomy (Prop. 3): iterative self-modification in agentic coding systems hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can only escape (run away) if that set expands. Frozen weights alone do not guarantee a bounded, safe regime.
-
Key practical directive: "Audit the scaffold, not the checkpoint." Since rewriting scaffolding (tools, verifiers, decomposition) expands what an agent reaches without touching a weight, frozen weights buy an eventual ceiling but no stationarity along the way (Cor. 4).
-
Sideways ceiling: Best-of- orchestration realizes the best worker's ceiling exactly (Prop. 5) — width buys rate, not budget. The weighted vote that could beat this ceiling requires diversity real workers lack: on 30 same-family workers, failure overlap sits at its maximum (), and a majority fails 23/55 (42%) of tasks (Prop. 6).
-
Empirical saturation: Per-round improvement decays toward zero on SWE-bench Lite (0.097 → 0.036 → 0.033 for Haiku; 0.156 → 0.068 → 0.029 for Sonnet), and churn decays geometrically across 401 production sessions — a shape shared with a pre-AI human baseline of 9,395 sessions.
-
Framework: The paper reads refinement as gradient boosting on the residual error between draft and target (a patch/git diff), then measures where that reading breaks: patches compose instead of standing beside each other for voting, and failures overlap. Both breaks are engineering choices, not laws about code, so they specify a harness worth building.
Introduction and Theoretical Foundation
Background and Motivation
Individual LLM calls are weak in a precise sense: they make errors and struggle with complex, multi-step reasoning. Agentic coding — deploying multiple agents that iteratively generate, critique, and refine code — addresses this, but the practice lacks rigorous mathematical foundation. Key motivating statistics:
- Anthropic reports >80% of code merged into its production codebase was authored by Claude as of May 2026 (up from low single digits before its coding-agent preview launched in February 2025).
- Engineers increasingly supply the goal rather than the method.
The paper asks: when does iterative self-modification run out of room, and is a frozen checkpoint evidence that it has?
Theoretical Foundation
The paper's central theoretical move is reading refinement as functional gradient descent (boosting):
Each agent refines an existing draft, predicting the residual between current code and target; it does not generate one from scratch, and that is what makes the process functional gradient descent.
This builds on:
- AdaBoost [Freund and Schapire, 1997] and margin-based generalization [Schapire et al., 1998]
- Boosting as functional gradient descent [Mason et al., 1999, Friedman, 2001]
- Structured prediction [Tsochantaridis et al., 2005]
- "Intelligence explosion" [Good, 1965, Bostrom, 2014] and the Gödel machine [Schmidhuber, 2003]
The paper distinguishes three regimes usually merged in the test-time-scaling literature:
- Search within a fixed class: inference compute searches within a class the weights and scaffold already fix.
- Test-time training: raises the ceiling itself (e.g., [Sun et al., 2020, Hardt and Sun, 2024, Akyürek et al., 2025]).
- Scaffold rewriting: between them — rewriting the scaffold with weights frozen.
The separation matters because "an auditor who checks only whether weights are frozen catches test-time training and misses scaffold expansion entirely."
Methodology
Formal Framework
Let be the space of specifications and the space of code artifacts (a draft , or the optimal target ), with a distribution over , and a loss (e.g., token/AST edit distance).
Definition 1 (Weak Coding Agent): A weak coding agent proposes a patch via a refinement operator , taking a draft and an edit to ; let indicate success on specification , iff the draft returned matches the target, . Every label is here, so .
The edge is on solving; Part B measures the edge on improving the previous draft:
on the pass fraction of task at round .
Key Theorems
Theorem 2 (Monotone Improvement under Refinement): Write for the draft after round at that round's best action , and suppose each satisfies the improvement property for quality function and . Then:
and since , the total improvement saturates: .
Proposition 3 (Stationarity Dichotomy): Consider the refinement game of Thm. 2.
- (i) Fixed ceiling ⇒ saturation: If for all , then , forcing : strict diminishing returns.
- (ii) Sustained edge ⇒ divergence: If the effective hypothesis class expands enough to sustain a uniform edge , then grows without bound and no fixed ceiling can hold.
Corollary 4 (Frozen-Weight Ceiling): Model a self-improving system as , frozen weights and mutable scaffold . Assume:
- (M1) class expandability: the effective hypothesis class can expand under a scaffold rewrite without altering ;
- (M2) ultimate ceiling: for all reachable by any scaffold the model can construct.
Then (i) the system saturates at a value no greater than ; (ii) escaping this ceiling requires raising itself — modifying weights or architecture — not only rewriting the scaffold.
Orchestration Results
Proposition 5 (Orchestration Selects a Ceiling; It Does Not Raise One): Best-of- selection attains exactly — an unconditional equality. Consulting all workers every round gives , yet since , the total improvement still saturates at : width buys rate, not budget.
Proposition 6 (Bounded Overlap Characterizes a Correct Vote): Let workers with outputs be combined by the unweighted majority vote, and let count the workers that fail point , with the tight overlap. Then the vote has zero training error iff for every — equivalently, iff . The condition is sharp: a single point at already makes the error nonzero.
Proposition 7 (Reused-Pool Dichotomy): Fix a pool of workers and run Algorithm 1 over it under any schedule at the optimal confidences. Let be the smallest zero-one training error attained by any nonnegative combination of the pool.
- (i) If , then , hence — the edge hypothesis provably fails by a finite, computable round.
- (ii) If instead the pool interpolates () with margin at coefficient sum , then for every distribution some member has edge , so the hypothesis survives every round.
Experimental Design
- Part A: Stylized theoretical validation — decision stumps on SWE-bench TF-IDF features with synthetic labels.
- Part B: Regime characterization — mini-swe-agent as the per-round weak learner, rounds, SWE-bench Lite tasks × 5 seeds, temperature 0.25, 2×3 design (feedback mode × model tier), 1,650 trajectories.
- Part C: Ecological evidence — 401 production sessions (1,211 commits, 23 repositories) vs. a 9,395-session pre-AI human baseline.
Empirical Validation / Results
Part B: SWE-bench Lite Results
Capability effect (large, real): Model tier explains roughly half the variance (, , partial ). Solve rates: Haiku 44–49% vs. Sonnet 85–89% (at of tests passing).
Feedback-mode effect (underpowered null): Neither feedback mode () nor the interaction () reaches significance. Power analysis: only 11%/17% power at this ; 80% power needs 7–9× the current sample.
Refinement saturates — mean per-round improvement on the running best:
| Tier | Round 1→2 | Round 2→3 | Round 3→4 |
|---|---|---|---|
| Haiku | 0.097 | 0.036 | 0.033 |
| Sonnet | 0.156 | 0.068 | 0.029 |
Improvement continues and shrinks toward zero — Thm. 2's saturation measured rather than argued.
Overlap measurement (30 workers = 6 arms × 5 seeds, 55 tasks):
- : some task defeats every worker.
- A majority fails 23/55 (42%) of tasks — the vote's error (0.418, 95% CI [0.291, 0.545]).
- The vote's error exceeds mean per-worker error 0.356, but the paired gap's interval includes zero (+0.062, 95% CI [−0.018, +0.142]).
Part C: Production Dynamics
Finding 1 — Churn converges: 70.0% of multi-commit sessions show decreasing churn across commits, at a median late/early churn ratio of 0.49 (95% CI [0.400, 0.793]). The ordering is genuinely temporal (commit-order permutation test, ).
Findings 2–3 — What that does and does not show: The pre-AI human baseline shows the same decay shape (geometric churn-decay factor against AI-tier ). The theory-specific discriminator returns a null: Opus shows no higher survivor ratio than Sonnet (, vs. , ; gap = −0.003, 95% CI [−0.075, +0.069], permutation ).
Theoretical and Practical Implications
Practitioner Lessons
- Spend on the first attempt, not the refinement loop — later rounds mostly leave the artifact untouched.
- Reach for capability before orchestration — orchestration cannot exceed the best worker's ceiling (Prop. 5), and the capability gap dominated every factor varied (partial ).
Governance Reading
The quantity to monitor is not whether weights are frozen but whether the scaffold is self-expanding.
A loop that only re-samples is bounded by construction; one that rewrites its tools or decomposition is not (M1, Cor. 4). Self-expansion doubles as a reward-hacking risk: a system licensed to rewrite its own tests may exploit rather than improve.
The Harness Specification
Algorithm 1 states what a loop would need for the transferred boosting guarantees to hold by construction:
- Retention of all proposals instead of an overwritten draft.
- An explicit distribution over specifications reweighted toward those still failing.
- Diversity sufficient to push the majority-failure rate below one half (exact condition: ).
Weighted voting over a shared instance is already boosting-exact; what deployed practice lacks is retention and diversity.
Open Problems
- Build the harness with retention, explicit distribution, and enforced : does a loop satisfying all three descend at the transferred rate?
- Run the vote across model families: does an ensemble from genuinely different vendors satisfy Prop. 6?
- Keep refining after the first pass: no unconditional per-round edge is estimable from data whose later rounds were never run.
- Hold weights fixed and vary the scaffold: the only design separating a genuinely fixed reachable class from ordinary editing economics.
Conclusion
Reading agentic coding as residual prediction yields a refinement game, a stationarity dichotomy (Prop. 3), and a frozen-weight ceiling. Key takeaways:
- A fixed reachable class and a bounded metric force strict diminishing returns, so freezing weights does not by itself guarantee a safe, saturating trajectory.
- The ceiling binds sideways too: best-of- orchestration realizes the best worker's ceiling exactly, the diversity a vote would need is absent on real workers, and re-consulting a fixed pool has a computable horizon unless that pool can already fit the work.
- The mapping is not boosting in the technical sense (patches compose rather than vote), but both remaining gaps are properties of harnesses as currently built rather than laws about code.
- What is measured is saturation — on SWE-bench and across 401 production sessions matched in shape by a pre-AI human baseline.
- What remains open is the mechanism: a genuinely fixed reachable class, or the ordinary economics of editing. Holding weights fixed while varying the scaffold would separate the two.
This data does not run that design; it argues for the experiment that would.
Key Definitions
| Term | Definition |
|---|---|
| Reachable edit set | The set of edits an agent can propose given its current scaffold and weights |
| Edge () | The margin over random guessing: |
| Tight overlap () | , the maximum number of workers failing any single point |
| Survivor ratio () | Net lines over total churn in a session |
| Scaffold | Everything mutable except weights: tools, verifiers, decomposition, prompts |
Related papers
- Merged, Not Measured: An Empirical Study of Performance Issues Fixed by Coding Agents
Agent performance fixes are merged based on the agent's track record and repository history, not the fix's content, tests, or measurements.
- WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
WideSWE, a new benchmark of 120 cross-repository tasks, shows top coding agents succeed only 42.5% of the time, revealing major gaps in multi-repo coordination.
- RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
RSIAgent, a training-free multi-agent framework, enables open-source models to outperform frontier closed-source models by autonomously exploring, verifying, and consolidating environment knowledge into reusable memory.