The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

Summary (Overview)

  • Core Problem: Self-improving LLM systems that keep changes scoring better on small evaluation sets suffer from selection bias—the "winner's curse"—where measured gains systematically overstate true held-out improvements.
  • Key Finding: With 16 selection items, final selection-set scores exceeded held-out accuracy by 13–20 points; with 256 items, this gap shrank to 1–5 points, confirming that small reused evaluation sets produce severely inflated performance estimates.
  • Theoretical Contribution: The paper formalizes keep-if-better loops as selection under measurement noise, deriving a Gaussian selection model with correlated candidate errors, yielding closed-form expressions for the winner's curse (Proposition 1) and a myopic Bayes acceptance rule (Proposition 2).
  • Empirical Results: Most proposals after the first rewrite are harmful (only 4–39% improve on their incumbent); tested acceptance rules (Bayes gates, select-then-confirm, McNemar tests) did not beat greedy acceptance over complete runs on reused selection sets.
  • Practical Recommendation: Self-improvement studies should report held-out gains with uncertainty; scoring the starting and current instruction on 64 fresh items removes average bias of reported gains (RMSE drops from 15.5 to 5.7 points).

Introduction and Theoretical Foundation

Background

Self-improving LLM systems (prompt optimizers, self-referential agents, LLM-guided search) share a common step: they propose a change, score it on a finite evaluation set, and keep it if it scores better than the current artifact. This selection score is also what appears in logs and validation outputs as "progress."

Motivation

Several recent audits suggest selection scores should not be taken at face value:

  • Evolved agent harnesses don't consistently beat test-time scaling at matched budget
  • Greedy acceptance commits many false or harmful edits
  • Measured self-training gains can be artifacts of the evaluation
  • A 41-paper methods audit found commonly omitted validity controls

Theoretical Foundation

The problem is framed as selection under measurement noise, building on:

  • Breeder's equation (Lush, 1937; Falconer and Mackay, 1996)
  • Optimizer's curse (Smith and Winkler, 2006; Harrison and March, 1984)
  • Selection-adjusted Bayesian inference (Dawid, 1994; Efron, 2011)
  • Online experimentation with adoption thresholds (Lee and Shen, 2018; Berman and Van den Bulte, 2022)

Key quantities defined:

  • Inflation: score on selection set minus held-out accuracy
  • Overstatement: gain measured on selection set minus held-out gain

Methodology

The Loop Model

An artifact cc (instruction, program, memory) conditions a solver with performance:

f(c)=Ex∼P[r(c,x)],r∈{0,1}f(c) = \mathbb{E}_{x \sim P}[r(c,x)], \quad r \in \{0,1\}

At generation tt, a proposer draws KK candidates from incumbent ctc_t, each scored on a selection set D=(x1,…,xn)D = (x_1, \ldots, x_n) paired with the incumbent:

Δ^k=1n∑i=1n[r(ck′,xi)−r(ct,xi)](1)\hat{\Delta}_k = \frac{1}{n}\sum_{i=1}^{n}\left[r(c_k', x_i) - r(c_t, x_i)\right] \tag{1}

Gaussian Selection Model

Assumption 1: Given history and fresh selection set:

  • Δk∼iidN(μ,s2)\Delta_k \stackrel{iid}{\sim} N(\mu, s^2) (true effects)
  • Δ^k=Δk+c+ηk\hat{\Delta}_k = \Delta_k + c + \eta_k with shared error c∼N(0,ec2)c \sim N(0, e_c^2) and idiosyncratic errors ηk∼iidN(0,eη2)\eta_k \stackrel{iid}{\sim} N(0, e_\eta^2)

Key parameters:

  • σw2=s2+eη2\sigma_w^2 = s^2 + e_\eta^2, hw2=s2/σw2h_w^2 = s^2/\sigma_w^2 (within-generation reliability)
  • h2=s2/(s2+e2)h^2 = s^2/(s^2 + e^2) (overall reliability)
  • κK=Emax⁡k≤KZk\kappa_K = \mathbb{E}\max_{k \leq K} Z_k (expected maximum of standard normals; κ4≈1.03\kappa_4 \approx 1.03, κ16≈1.77\kappa_{16} \approx 1.77)

Proposition 1 (Selection Differential and Overstatement)

Under Assumption 1, the measured gain of the best candidate k^\hat{k} overstates its true gain by:

E[Δ^k^−Δk^]=(1−hw2)σwκK+E[c](2)\mathbb{E}[\hat{\Delta}_{\hat{k}} - \Delta_{\hat{k}}] = (1 - h_w^2)\sigma_w \kappa_K + \mathbb{E}[c] \tag{2}

This is the optimizer's curse—regression to the mean of the selected candidate.

Proposition 2 (Myopic Bayes Commit)

The rule maximizing expected true improvement commits the candidate with the largest posterior mean:

E[Δk∣Δ^]=μ+hw2(Δ^k−mˉ)+hw2σw2σw2+Kec2(mˉ−μ)(3)\mathbb{E}[\Delta_k \mid \hat{\Delta}] = \mu + h_w^2(\hat{\Delta}_k - \bar{m}) + \frac{h_w^2 \sigma_w^2}{\sigma_w^2 + K e_c^2}(\bar{m} - \mu) \tag{3}

With independent noise (ec=0e_c = 0) and s>0s > 0, this reduces to a threshold: Δ^k^>τ∗=−μe2/s2\hat{\Delta}_{\hat{k}} > \tau^* = -\mu e^2/s^2.

Proposition 3 (Goodhart Regime)

With s=0s = 0, μ<0\mu < 0, and fresh selection sets each generation, greedy commits with probability p=Pr[c+eηmax⁡kZk>∣μ∣]p = \text{Pr}[c + e_\eta \max_k Z_k > |\mu|] per generation; held-out performance drifts down by ∣μ∣p|\mu|p per generation while every commit reports positive gain.

Acceptance Rules Tested

  • Greedy: commit if Δ^k^>0\hat{\Delta}_{\hat{k}} > 0
  • McNemar: one-sided exact test against incumbent at level α\alpha
  • PACE: anytime-valid test (Shawn, 2026)
  • Select-then-confirm: choose on half of DD, commit if winner also beats incumbent on other half
  • Bayes gates: apply Equation (3) with regularized noise terms

Experimental Setup

  • Models: Qwen2.5-1.5B/7B-Instruct (exploratory/confirmatory), Qwen3.5-4B (current model), Qwen3.8-27B (proposer for GEPA)
  • Tasks: TREC-50, Banking77, GSM8K
  • 600-item gold (held-out) set per task, excluded from selection
  • Selection sets of n∈{16,64,256}n \in \{16, 64, 256\} items, reused across generations
  • K=4K = 4 candidates per generation, T=20T = 20 generations

Empirical Validation / Results

5.1 Proposals Are Mostly Harmful

  • First rewrite of a one-sentence seed often improves (75–100% on TREC)
  • After that, only 4–39% of rewrites improve on their incumbent
  • Mean held-out effects between −0.2 and −13.5 points
  • Candidates resemble each other more than the incumbent (vcc≈0.44–0.60vv_{cc} \approx 0.44–0.60v)

The noise model's predictions match observed ratios of held-out to measured selection differentials (mean absolute error 0.034 with empirical distribution, 0.064 with Gaussian approximation). At n=16n = 16, the winner's held-out selection differential is only 1–51% of its measured one.

5.2 Commits on Reused Selection Sets

For 72 first-generation commits, observed vs. model-predicted overstatement:

nnObservedModel
169.1 pts9.0 pts
643.2 pts3.8 pts
2560.1 pts1.3 pts

Lock-in effect: commits against a selected incumbent overstate less at every nn (5.2 vs. 9.7, 1.9 vs. 3.7, 0.5 vs. 1.1 points), consistent with the incumbent's positive selection error acting as an implicit threshold.

5.3 Pre-registered Confirmatory Study

Table 1: Pre-registered tests (points; mean ± s.e. of seed-paired differences; Holm-adjusted p-values)

SettingDifference (pts)Wins/Pairsp
H1: gain(n=256) − gain(n=16), greedy
1.5B TREC+10.5±2.67/80.010
7B TREC+6.4±0.78/8<0.001
1.5B GSM8K+1.6±2.24/80.483
H2: gap(n=16) − gap(n=256), greedy
1.5B TREC+15.9±2.78/80.002
1.5B Banking77+10.5±4.67/80.058
7B TREC+15.3±4.77/80.037
1.5B GSM8K+14.7±4.48/80.037
H4: pilot gate − greedy at n=16
Pooled, 32 pairs−2.1±1.110/320.062
1.5B GSM8K−5.7±1.61/80.032

Key results:

  • H1 (partial support): Held-out gains rose with nn on TREC (13.7→24.2 pts for 1.5B; 14.0→20.4 for 7B) but not GSM8K
  • H2 (supported): Final proxy exceeded held-out accuracy by 13–20 pts (n=16) vs. 1–5 pts (n=256)
  • H3 (not replicated): Goodhart loss on Banking77 did not replicate (−0.6±0.7 pts)
  • H4/H5 (not supported): Bayes gates and select-then-confirm did not beat greedy

5.4 Current Models, Strong Starts, and Optimizers

Qwen3.5-4B self-improvement from competent start:

  • Only 5% of TREC rewrites improve (mean effect −14 pts)
  • With 16 items: reported gains of 12.5–14.1 pts, held-out changes of −2.2 to 0.0
  • Overstatement fell from 16.3 to 1.4 pts (TREC) and 11.5 to 1.5 (Banking77) when moving from 16 to 256 items

GEPA optimizer (16 vs. 256 validation items, 1000 metric calls):

  • TREC: reported 38.8 pts vs. real 19.2 pts (16 items); ≤3 pts difference with 256 items
  • Banking77: reported 18.8 pts vs. real 1.2±1.7 pts (16 items)
  • GSM8K: reported 6.2 pts for a program 0.9 pts worse

MIPROv2:

  • TREC: 11.4 held-out pts with 256 items vs. 5.8 with 16 (paired diff 5.6±0.75.6 \pm 0.7, p=0.001p = 0.001)
  • Banking77: 16-item arm returned programs 1.6 pts worse than seed, though validation implied 5-pt gain

Theoretical and Practical Implications

Theoretical Contributions

  1. Unified selection model: Extends classical selection theory (breeder's equation, optimizer's curse) to LLM self-improvement with correlated candidate errors and reused evaluation sets
  2. Lock-in mechanism: Reused selection sets create an implicit threshold via the incumbent's selection error—a form of regressional Goodhart
  3. Bayes acceptance rule: Provides calibrated thresholds for when to commit a candidate, though optimality is only myopic

Practical Implications

  1. Reporting standards: Self-improvement studies should report held-out gains with standard errors; a 64-item audit set (128 evaluations, ~10% of nominal budget) removes average bias of reported gains (RMSE: 15.5 → 5.7 pts)
  2. Selection set size matters: Larger selection sets reduce inflation but may not improve held-out gains (TREC: yes; GSM8K: no; GEPA: no, because smaller sets allow more search)
  3. Acceptance rules are not a panacea: Calibrated thresholds help isolated decisions but did not improve complete runs on reused sets
  4. Proposal quality dominates: The proposer's ability to generate genuinely better candidates matters more than the acceptance rule

Conclusion

Main Takeaways

  • The winner's curse is real and quantifiable in LLM self-improvement loops, with the Gaussian selection model accurately predicting its size on average
  • Small reused evaluation sets produce severely inflated self-reported gains (13–20 points with 16 items vs. 1–5 with 256)
  • Most proposals are harmful after the first rewrite; the system's "improvements" are largely selection artifacts
  • No tested acceptance rule reliably beats greedy on complete runs with reused selection sets

Future Directions

  • Modeling the dynamics of lock-in more precisely
  • Testing on agentic and code tasks with stochastic rollouts
  • Developing online estimates of proposal bias (μ\mu) without separate pilots
  • Investigating whether fresh selection sets per generation can avoid accumulated inflation

Limitations

  • Artifacts limited to instructions; deterministic scoring rules
  • Largest proposer has 27B parameters
  • Bayes gates require a pilot with a large evaluation set
  • No claims about recursive self-improvement or long-horizon dynamics

Related papers