# The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

> Self-improving LLM loops on small reused evaluation sets inflate reported gains by 13–20 points, with most proposals after the first rewrite being harmful.

- **Source:** [arXiv](https://arxiv.org/abs/2610.09239)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/YBi2Bm
- **Whiteboard:** https://picx.dev/p/YBi2Bm/image

## Summary

# The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules

## Summary (Overview)

- **Core Problem**: Self-improving LLM systems that keep changes scoring better on small evaluation sets suffer from **selection bias**—the "winner's curse"—where measured gains systematically overstate true held-out improvements.
- **Key Finding**: With 16 selection items, final selection-set scores exceeded held-out accuracy by **13–20 points**; with 256 items, this gap shrank to **1–5 points**, confirming that small reused evaluation sets produce severely inflated performance estimates.
- **Theoretical Contribution**: The paper formalizes keep-if-better loops as selection under measurement noise, deriving a Gaussian selection model with correlated candidate errors, yielding closed-form expressions for the winner's curse (Proposition 1) and a myopic Bayes acceptance rule (Proposition 2).
- **Empirical Results**: Most proposals after the first rewrite are harmful (only 4–39% improve on their incumbent); tested acceptance rules (Bayes gates, select-then-confirm, McNemar tests) did **not** beat greedy acceptance over complete runs on reused selection sets.
- **Practical Recommendation**: Self-improvement studies should report held-out gains with uncertainty; scoring the starting and current instruction on 64 fresh items removes average bias of reported gains (RMSE drops from 15.5 to 5.7 points).

---

## Introduction and Theoretical Foundation

### Background

Self-improving LLM systems (prompt optimizers, self-referential agents, LLM-guided search) share a common step: they **propose a change, score it on a finite evaluation set, and keep it if it scores better** than the current artifact. This selection score is also what appears in logs and validation outputs as "progress."

### Motivation

Several recent audits suggest selection scores should not be taken at face value:
- Evolved agent harnesses don't consistently beat test-time scaling at matched budget
- Greedy acceptance commits many false or harmful edits
- Measured self-training gains can be artifacts of the evaluation
- A 41-paper methods audit found commonly omitted validity controls

### Theoretical Foundation

The problem is framed as **selection under measurement noise**, building on:
- **Breeder's equation** (Lush, 1937; Falconer and Mackay, 1996)
- **Optimizer's curse** (Smith and Winkler, 2006; Harrison and March, 1984)
- **Selection-adjusted Bayesian inference** (Dawid, 1994; Efron, 2011)
- **Online experimentation** with adoption thresholds (Lee and Shen, 2018; Berman and Van den Bulte, 2022)

Key quantities defined:
- **Inflation**: score on selection set minus held-out accuracy
- **Overstatement**: gain measured on selection set minus held-out gain

---

## Methodology

### The Loop Model

An artifact $c$ (instruction, program, memory) conditions a solver with performance:

$$f(c) = \mathbb{E}_{x \sim P}[r(c,x)], \quad r \in \{0,1\}$$

At generation $t$, a proposer draws $K$ candidates from incumbent $c_t$, each scored on a selection set $D = (x_1, \ldots, x_n)$ paired with the incumbent:

$$\hat{\Delta}_k = \frac{1}{n}\sum_{i=1}^{n}\left[r(c_k', x_i) - r(c_t, x_i)\right] \tag{1}$$

### Gaussian Selection Model

**Assumption 1**: Given history and fresh selection set:
- $\Delta_k \stackrel{iid}{\sim} N(\mu, s^2)$ (true effects)
- $\hat{\Delta}_k = \Delta_k + c + \eta_k$ with shared error $c \sim N(0, e_c^2)$ and idiosyncratic errors $\eta_k \stackrel{iid}{\sim} N(0, e_\eta^2)$

Key parameters:
- $\sigma_w^2 = s^2 + e_\eta^2$, $h_w^2 = s^2/\sigma_w^2$ (within-generation reliability)
- $h^2 = s^2/(s^2 + e^2)$ (overall reliability)
- $\kappa_K = \mathbb{E}\max_{k \leq K} Z_k$ (expected maximum of standard normals; $\kappa_4 \approx 1.03$, $\kappa_{16} \approx 1.77$)

### Proposition 1 (Selection Differential and Overstatement)

Under Assumption 1, the measured gain of the best candidate $\hat{k}$ overstates its true gain by:

$$\mathbb{E}[\hat{\Delta}_{\hat{k}} - \Delta_{\hat{k}}] = (1 - h_w^2)\sigma_w \kappa_K + \mathbb{E}[c] \tag{2}$$

This is the **optimizer's curse**—regression to the mean of the selected candidate.

### Proposition 2 (Myopic Bayes Commit)

The rule maximizing expected true improvement commits the candidate with the largest posterior mean:

$$\mathbb{E}[\Delta_k \mid \hat{\Delta}] = \mu + h_w^2(\hat{\Delta}_k - \bar{m}) + \frac{h_w^2 \sigma_w^2}{\sigma_w^2 + K e_c^2}(\bar{m} - \mu) \tag{3}$$

With independent noise ($e_c = 0$) and $s > 0$, this reduces to a threshold: $\hat{\Delta}_{\hat{k}} > \tau^* = -\mu e^2/s^2$.

### Proposition 3 (Goodhart Regime)

With $s = 0$, $\mu < 0$, and fresh selection sets each generation, greedy commits with probability $p = \text{Pr}[c + e_\eta \max_k Z_k > |\mu|]$ per generation; held-out performance drifts down by $|\mu|p$ per generation while every commit reports positive gain.

### Acceptance Rules Tested

- **Greedy**: commit if $\hat{\Delta}_{\hat{k}} > 0$
- **McNemar**: one-sided exact test against incumbent at level $\alpha$
- **PACE**: anytime-valid test (Shawn, 2026)
- **Select-then-confirm**: choose on half of $D$, commit if winner also beats incumbent on other half
- **Bayes gates**: apply Equation (3) with regularized noise terms

### Experimental Setup

- Models: Qwen2.5-1.5B/7B-Instruct (exploratory/confirmatory), Qwen3.5-4B (current model), Qwen3.8-27B (proposer for GEPA)
- Tasks: TREC-50, Banking77, GSM8K
- 600-item gold (held-out) set per task, excluded from selection
- Selection sets of $n \in \{16, 64, 256\}$ items, reused across generations
- $K = 4$ candidates per generation, $T = 20$ generations

---

## Empirical Validation / Results

### 5.1 Proposals Are Mostly Harmful

- First rewrite of a one-sentence seed often improves (75–100% on TREC)
- After that, only **4–39%** of rewrites improve on their incumbent
- Mean held-out effects between −0.2 and −13.5 points
- Candidates resemble each other more than the incumbent ($v_{cc} \approx 0.44–0.60v$)

The noise model's predictions match observed ratios of held-out to measured selection differentials (mean absolute error 0.034 with empirical distribution, 0.064 with Gaussian approximation). At $n = 16$, the winner's held-out selection differential is only **1–51%** of its measured one.

### 5.2 Commits on Reused Selection Sets

For 72 first-generation commits, observed vs. model-predicted overstatement:

| $n$ | Observed | Model |
|-----|----------|-------|
| 16  | 9.1 pts  | 9.0 pts |
| 64  | 3.2 pts  | 3.8 pts |
| 256 | 0.1 pts  | 1.3 pts |

**Lock-in effect**: commits against a selected incumbent overstate less at every $n$ (5.2 vs. 9.7, 1.9 vs. 3.7, 0.5 vs. 1.1 points), consistent with the incumbent's positive selection error acting as an implicit threshold.

### 5.3 Pre-registered Confirmatory Study

**Table 1: Pre-registered tests** (points; mean ± s.e. of seed-paired differences; Holm-adjusted p-values)

| Setting | Difference (pts) | Wins/Pairs | p |
|---------|-----------------|------------|---|
| **H1: gain(n=256) − gain(n=16), greedy** | | | |
| 1.5B TREC | +10.5±2.6 | 7/8 | 0.010 |
| 7B TREC | +6.4±0.7 | 8/8 | <0.001 |
| 1.5B GSM8K | +1.6±2.2 | 4/8 | 0.483 |
| **H2: gap(n=16) − gap(n=256), greedy** | | | |
| 1.5B TREC | +15.9±2.7 | 8/8 | 0.002 |
| 1.5B Banking77 | +10.5±4.6 | 7/8 | 0.058 |
| 7B TREC | +15.3±4.7 | 7/8 | 0.037 |
| 1.5B GSM8K | +14.7±4.4 | 8/8 | 0.037 |
| **H4: pilot gate − greedy at n=16** | | | |
| Pooled, 32 pairs | −2.1±1.1 | 10/32 | 0.062 |
| 1.5B GSM8K | −5.7±1.6 | 1/8 | 0.032 |

**Key results**:
- **H1 (partial support)**: Held-out gains rose with $n$ on TREC (13.7→24.2 pts for 1.5B; 14.0→20.4 for 7B) but not GSM8K
- **H2 (supported)**: Final proxy exceeded held-out accuracy by 13–20 pts (n=16) vs. 1–5 pts (n=256)
- **H3 (not replicated)**: Goodhart loss on Banking77 did not replicate (−0.6±0.7 pts)
- **H4/H5 (not supported)**: Bayes gates and select-then-confirm did not beat greedy

### 5.4 Current Models, Strong Starts, and Optimizers

**Qwen3.5-4B self-improvement from competent start**:
- Only 5% of TREC rewrites improve (mean effect −14 pts)
- With 16 items: reported gains of 12.5–14.1 pts, held-out changes of −2.2 to 0.0
- Overstatement fell from 16.3 to 1.4 pts (TREC) and 11.5 to 1.5 (Banking77) when moving from 16 to 256 items

**GEPA optimizer** (16 vs. 256 validation items, 1000 metric calls):
- TREC: reported 38.8 pts vs. real 19.2 pts (16 items); ≤3 pts difference with 256 items
- Banking77: reported 18.8 pts vs. real 1.2±1.7 pts (16 items)
- GSM8K: reported 6.2 pts for a program 0.9 pts *worse*

**MIPROv2**:
- TREC: 11.4 held-out pts with 256 items vs. 5.8 with 16 (paired diff $5.6 \pm 0.7$, $p = 0.001$)
- Banking77: 16-item arm returned programs 1.6 pts worse than seed, though validation implied 5-pt gain

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Unified selection model**: Extends classical selection theory (breeder's equation, optimizer's curse) to LLM self-improvement with correlated candidate errors and reused evaluation sets
2. **Lock-in mechanism**: Reused selection sets create an implicit threshold via the incumbent's selection error—a form of regressional Goodhart
3. **Bayes acceptance rule**: Provides calibrated thresholds for when to commit a candidate, though optimality is only myopic

### Practical Implications

1. **Reporting standards**: Self-improvement studies should report held-out gains with standard errors; a 64-item audit set (128 evaluations, ~10% of nominal budget) removes average bias of reported gains (RMSE: 15.5 → 5.7 pts)
2. **Selection set size matters**: Larger selection sets reduce inflation but may not improve held-out gains (TREC: yes; GSM8K: no; GEPA: no, because smaller sets allow more search)
3. **Acceptance rules are not a panacea**: Calibrated thresholds help isolated decisions but did not improve complete runs on reused sets
4. **Proposal quality dominates**: The proposer's ability to generate genuinely better candidates matters more than the acceptance rule

---

## Conclusion

### Main Takeaways

- **The winner's curse is real and quantifiable** in LLM self-improvement loops, with the Gaussian selection model accurately predicting its size on average
- **Small reused evaluation sets produce severely inflated self-reported gains** (13–20 points with 16 items vs. 1–5 with 256)
- **Most proposals are harmful** after the first rewrite; the system's "improvements" are largely selection artifacts
- **No tested acceptance rule reliably beats greedy** on complete runs with reused selection sets

### Future Directions

- Modeling the dynamics of lock-in more precisely
- Testing on agentic and code tasks with stochastic rollouts
- Developing online estimates of proposal bias ($\mu$) without separate pilots
- Investigating whether fresh selection sets per generation can avoid accumulated inflation

### Limitations

- Artifacts limited to instructions; deterministic scoring rules
- Largest proposer has 27B parameters
- Bayes gates require a pilot with a large evaluation set
- No claims about recursive self-improvement or long-horizon dynamics

---

_Markdown view of https://picx.dev/p/YBi2Bm, served by PicX — AI-generated visual whiteboard summaries of research papers._
