# Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding

> Frozen weights do not guarantee safe saturation in agentic coding; auditors must monitor scaffold expansion, not just checkpoint freezing.

- **Source:** [arXiv](https://arxiv.org/abs/2609.34924)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/rUmGhy
- **Whiteboard:** https://picx.dev/p/rUmGhy/image

## Summary

# Summary of "Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding"

## Summary (Overview)

- **Core theoretical contribution**: The paper establishes a **stationarity dichotomy** (Prop. 3): iterative self-modification in agentic coding systems hits strict diminishing returns whenever the agent's *reachable set of edits* stays fixed, and can only escape (run away) if that set expands. Frozen weights alone do **not** guarantee a bounded, safe regime.

- **Key practical directive**: **"Audit the scaffold, not the checkpoint."** Since rewriting scaffolding (tools, verifiers, decomposition) expands what an agent reaches *without touching a weight*, frozen weights buy an eventual ceiling but no stationarity along the way (Cor. 4).

- **Sideways ceiling**: Best-of-$k$ orchestration realizes the best worker's ceiling **exactly** (Prop. 5) — width buys rate, not budget. The weighted vote that could beat this ceiling requires diversity real workers lack: on 30 same-family workers, failure overlap sits at its maximum ($m^\star = k$), and a majority fails 23/55 (42%) of tasks (Prop. 6).

- **Empirical saturation**: Per-round improvement decays toward zero on SWE-bench Lite (0.097 → 0.036 → 0.033 for Haiku; 0.156 → 0.068 → 0.029 for Sonnet), and churn decays geometrically across 401 production sessions — a shape shared with a pre-AI human baseline of 9,395 sessions.

- **Framework**: The paper reads refinement as **gradient boosting on the residual error** between draft and target (a patch/git diff), then measures where that reading breaks: patches *compose* instead of standing beside each other for voting, and failures overlap. Both breaks are engineering choices, not laws about code, so they specify a harness worth building.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Individual LLM calls are weak in a precise sense: they make errors and struggle with complex, multi-step reasoning. Agentic coding — deploying multiple agents that iteratively generate, critique, and refine code — addresses this, but the practice lacks rigorous mathematical foundation. Key motivating statistics:

- Anthropic reports >80% of code merged into its production codebase was authored by Claude as of May 2026 (up from low single digits before its coding-agent preview launched in February 2025).
- Engineers increasingly supply the goal rather than the method.

The paper asks: **when does iterative self-modification run out of room, and is a frozen checkpoint evidence that it has?**

### Theoretical Foundation

The paper's central theoretical move is reading refinement as **functional gradient descent** (boosting):

> Each agent refines an existing draft, predicting the residual between current code and target; it does not generate one from scratch, and that is what makes the process functional gradient descent.

This builds on:
- **AdaBoost** [Freund and Schapire, 1997] and margin-based generalization [Schapire et al., 1998]
- Boosting as functional gradient descent [Mason et al., 1999, Friedman, 2001]
- Structured prediction [Tsochantaridis et al., 2005]
- "Intelligence explosion" [Good, 1965, Bostrom, 2014] and the Gödel machine [Schmidhuber, 2003]

The paper distinguishes **three regimes** usually merged in the test-time-scaling literature:

1. **Search within a fixed class**: inference compute searches within a class the weights and scaffold already fix.
2. **Test-time training**: raises the ceiling itself (e.g., [Sun et al., 2020, Hardt and Sun, 2024, Akyürek et al., 2025]).
3. **Scaffold rewriting**: between them — rewriting the scaffold with weights frozen.

The separation matters because "an auditor who checks only whether weights are frozen catches test-time training and misses scaffold expansion entirely."

---

## Methodology

### Formal Framework

Let $X$ be the space of specifications and $Y$ the space of code artifacts (a draft $z$, or the optimal target $y$), with a distribution $D$ over $X \times Y$, and a loss $L: Y \times Y \to \mathbb{R}_{\geq 0}$ (e.g., token/AST edit distance).

**Definition 1 (Weak Coding Agent)**: A weak coding agent proposes a patch via a refinement operator $R_t: Y \times A \to Y$, taking a draft $z$ and an edit $a \in A$ to $z' = R_t(z, a)$; let $h_t(x) \in \{-1, +1\}$ indicate success on specification $x$, $h_t(x) = +1$ iff the draft returned matches the target, $L(y, z') = 0$. Every label is $+1$ here, so $\text{err}_D(h_t) = \Pr[h_t(x) = -1] \leq \frac{1}{2} - \gamma_t$.

The edge $\gamma_t$ is on *solving*; Part B measures the edge on *improving* the previous draft:
$$\tilde{\gamma}_t = \frac{1}{N} \sum_i \mathbb{1}[q_{i,t} > q_{i,t-1}] - \frac{1}{2}$$
on the pass fraction $q_{i,t}$ of task $i$ at round $t$.

### Key Theorems

**Theorem 2 (Monotone Improvement under Refinement)**: Write $z_t = R_t(z_{t-1}, a_t^*)$ for the draft after round $t$ at that round's best action $a_t^*$, and suppose each $R_t$ satisfies the improvement property $\mathbb{E}_{y \sim D(\cdot|x)}[V(R_t(z, a_t^*))] \geq V(z) + \eta_t$ for quality function $V: Y \to [0,1]$ and $\eta_t > 0$. Then:
$$\mathbb{E}[V(z_T)] \geq V(z_0) + \sum_{t=1}^{T} \eta_t$$
and since $V \leq 1$, the total improvement saturates: $\sum_{t=1}^{\infty} \eta_t \leq 1 - V(z_0)$.

**Proposition 3 (Stationarity Dichotomy)**: Consider the refinement game of Thm. 2.

- **(i) Fixed ceiling ⇒ saturation**: If $V(z) \leq B$ for all $z$, then $\sum_{t=1}^{\infty} \eta_t \leq B - \mathbb{E}[V(z_0)]$, forcing $\eta_t \to 0$: strict diminishing returns.
- **(ii) Sustained edge ⇒ divergence**: If the effective hypothesis class expands enough to sustain a uniform edge $\eta_t \geq c > 0$, then $\mathbb{E}[V(z_T)]$ grows without bound and no fixed ceiling can hold.

**Corollary 4 (Frozen-Weight Ceiling)**: Model a self-improving system as $S = (W, C)$, frozen weights $W$ and mutable scaffold $C$. Assume:
- **(M1) class expandability**: the effective hypothesis class $H_t(C)$ can expand under a scaffold rewrite without altering $W$;
- **(M2) ultimate ceiling**: $V(z) \leq V^*(W)$ for all $z$ reachable by any scaffold the model can construct.

Then (i) the system saturates at a value no greater than $V^*(W)$; (ii) escaping this ceiling requires raising $V^*(W)$ itself — modifying weights or architecture — not only rewriting the scaffold.

### Orchestration Results

**Proposition 5 (Orchestration Selects a Ceiling; It Does Not Raise One)**: Best-of-$k$ selection attains $V^*_{\text{orch}} = \max_i V^*(W, C_i)$ exactly — an unconditional equality. Consulting all $k$ workers every round gives $\mathbb{E}[V(z_T)] \geq V(z_0) + \sum_{t=1}^{T} \max_i \eta_t^{(i)}$, yet since $V \leq 1$, the total improvement still saturates at $1 - V(z_0)$: **width buys rate, not budget**.

**Proposition 6 (Bounded Overlap Characterizes a Correct Vote)**: Let $k$ workers with $\pm 1$ outputs be combined by the unweighted majority vote, and let $m_i := \#\{t < k : y_i h_t(x_i) = -1\}$ count the workers that fail point $i$, with $m^\star := \max_i m_i$ the tight overlap. Then the vote has zero training error iff $2m_i < k$ for every $i$ — equivalently, iff $m^\star < k/2$. The condition is sharp: a single point at $2m_i = k$ already makes the error nonzero.

**Proposition 7 (Reused-Pool Dichotomy)**: Fix a pool of $k$ workers and run Algorithm 1 over it under any schedule at the optimal confidences. Let $\epsilon_0$ be the smallest zero-one training error attained by any nonnegative combination of the pool.

- (i) If $\epsilon_0 > 0$, then $\epsilon_0 \leq e^{-2\gamma^2 T}$, hence $T \leq \log(1/\epsilon_0) / 2\gamma^2$ — the edge hypothesis provably fails by a finite, computable round.
- (ii) If instead the pool interpolates ($\epsilon_0 = 0$) with margin $\theta$ at coefficient sum $S$, then for every distribution some member has edge $\gamma = \theta/2S > 0$, so the hypothesis survives every round.

### Experimental Design

- **Part A**: Stylized theoretical validation — decision stumps on SWE-bench TF-IDF features with synthetic labels.
- **Part B**: Regime characterization — mini-swe-agent as the per-round weak learner, $T = 4$ rounds, $K = 55$ SWE-bench Lite tasks × 5 seeds, temperature 0.25, 2×3 design (feedback mode × model tier), 1,650 trajectories.
- **Part C**: Ecological evidence — 401 production sessions (1,211 commits, 23 repositories) vs. a 9,395-session pre-AI human baseline.

---

## Empirical Validation / Results

### Part B: SWE-bench Lite Results

**Capability effect (large, real)**: Model tier explains roughly half the variance ($F(1, 54) = 53.46$, $p < 0.0001$, partial $\eta^2 = 0.497$). Solve rates: Haiku 44–49% vs. Sonnet 85–89% (at $\tau \geq 0.6$ of tests passing).

**Feedback-mode effect (underpowered null)**: Neither feedback mode ($p = 0.61$) nor the interaction ($p = 0.45$) reaches significance. Power analysis: only 11%/17% power at this $K$; 80% power needs 7–9× the current sample.

**Refinement saturates** — mean per-round improvement on the running best:

| Tier | Round 1→2 | Round 2→3 | Round 3→4 |
|------|-----------|-----------|-----------|
| Haiku | 0.097 | 0.036 | 0.033 |
| Sonnet | 0.156 | 0.068 | 0.029 |

Improvement continues and shrinks toward zero — Thm. 2's saturation measured rather than argued.

**Overlap measurement** (30 workers = 6 arms × 5 seeds, 55 tasks):

- $m^\star = 30 = k$: some task defeats every worker.
- A majority fails 23/55 (42%) of tasks — the vote's error (0.418, 95% CI [0.291, 0.545]).
- The vote's error exceeds mean per-worker error 0.356, but the paired gap's interval includes zero (+0.062, 95% CI [−0.018, +0.142]).

### Part C: Production Dynamics

**Finding 1 — Churn converges**: 70.0% of multi-commit sessions show decreasing churn across commits, at a median late/early churn ratio of 0.49 (95% CI [0.400, 0.793]). The ordering is genuinely temporal (commit-order permutation test, $p < 0.0001$).

**Findings 2–3 — What that does and does not show**: The pre-AI human baseline shows the same decay shape (geometric churn-decay factor $\rho_{\text{human}} = 0.86$ against AI-tier $\rho \approx 0.77$). The theory-specific discriminator returns a null: Opus shows no higher survivor ratio than Sonnet ($\psi_{\text{Opus}} \approx 0.55$, $n = 158$ vs. $\psi_{\text{Sonnet}} \approx 0.55$, $n = 177$; gap = −0.003, 95% CI [−0.075, +0.069], permutation $p = 0.93$).

---

## Theoretical and Practical Implications

### Practitioner Lessons

1. **Spend on the first attempt, not the refinement loop** — later rounds mostly leave the artifact untouched.
2. **Reach for capability before orchestration** — orchestration cannot exceed the best worker's ceiling (Prop. 5), and the capability gap dominated every factor varied (partial $\eta^2 = 0.497$).

### Governance Reading

> The quantity to monitor is not whether weights are frozen but whether the scaffold is self-expanding.

A loop that only re-samples is bounded by construction; one that rewrites its tools or decomposition is not (M1, Cor. 4). Self-expansion doubles as a reward-hacking risk: a system licensed to rewrite its own tests may exploit $L(Y, \hat{Y})$ rather than improve.

### The Harness Specification

Algorithm 1 states what a loop would need for the transferred boosting guarantees to hold by construction:

1. **Retention** of all $T$ proposals instead of an overwritten draft.
2. An explicit **distribution $D_t$ over specifications** reweighted toward those still failing.
3. **Diversity** sufficient to push the majority-failure rate below one half (exact condition: $m^\star < k/2$).

Weighted voting over a shared instance is already boosting-exact; what deployed practice lacks is retention and diversity.

### Open Problems

1. **Build the harness** with retention, explicit distribution, and enforced $m^\star < k/2$: does a loop satisfying all three descend at the transferred rate?
2. **Run the vote across model families**: does an ensemble from genuinely different vendors satisfy Prop. 6?
3. **Keep refining after the first pass**: no unconditional per-round edge is estimable from data whose later rounds were never run.
4. **Hold weights fixed and vary the scaffold**: the only design separating a genuinely fixed reachable class from ordinary editing economics.

---

## Conclusion

Reading agentic coding as residual prediction yields a refinement game, a stationarity dichotomy (Prop. 3), and a frozen-weight ceiling. Key takeaways:

- A fixed reachable class and a bounded metric force strict diminishing returns, so **freezing weights does not by itself guarantee a safe, saturating trajectory**.
- The ceiling binds sideways too: best-of-$k$ orchestration realizes the best worker's ceiling exactly, the diversity a vote would need is absent on real workers, and re-consulting a fixed pool has a computable horizon unless that pool can already fit the work.
- The mapping is not boosting in the technical sense (patches compose rather than vote), but both remaining gaps are properties of harnesses as currently built rather than laws about code.
- **What is measured is saturation** — on SWE-bench and across 401 production sessions matched in shape by a pre-AI human baseline.
- **What remains open is the mechanism**: a genuinely fixed reachable class, or the ordinary economics of editing. Holding weights fixed while varying the scaffold would separate the two.

> This data does not run that design; it argues for the experiment that would.

---

## Key Definitions

| Term | Definition |
|------|------------|
| **Reachable edit set** | The set of edits an agent can propose given its current scaffold and weights |
| **Edge ($\gamma_t$)** | The margin over random guessing: $\text{err}_D(h_t) \leq \frac{1}{2} - \gamma_t$ |
| **Tight overlap ($m^\star$)** | $\max_i m_i$, the maximum number of workers failing any single point |
| **Survivor ratio ($\psi$)** | Net lines over total churn in a session |
| **Scaffold** | Everything mutable except weights: tools, verifiers, decomposition, prompts |

---

_Markdown view of https://picx.dev/p/rUmGhy, served by PicX — AI-generated visual whiteboard summaries of research papers._
