# VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

> VERSE shows LLM optimizers improve by evolving their own harness, but only when execution-based verification tools are provided, boosting SWE-rebench accuracy across all baselines.

- **Source:** [arXiv](https://arxiv.org/abs/2610.02616)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/Gv8uGC
- **Whiteboard:** https://picx.dev/p/Gv8uGC/image

## Summary

## Summary (Overview)

- **VERSE** (Verified Self-Evolving optimizer for agent harnesses) introduces a framework where an LLM optimizer can evolve both the executor's harness (prompts, tools, workflow) *and* its own optimizer harness, while keeping model weights fixed.
- Two key motivating observations: (1) optimizer self-evolution only improves performance when execution-based verification is available; (2) self-evolving optimizers naturally build tools for failure attribution, verification, training audits, and workflow control.
- VERSE provides four components: attribution (trace minimization), verification tools (verification, replay, perturbation), training audit, and optimizer self-evolution.
- The paper includes a theoretical framework viewing harness evolution as bilevel learning, proving that execution checks can lower the minimum rate of wrong repairs when evidence is ambiguous.
- Experiments on SWE-rebench show VERSE improves all four evaluated harness optimizers on held-out tasks and out-of-distribution tasks across five languages, achieving 42.3% (in-distribution) and 37.7% (OOD) accuracy vs. 39.2% and 29.3% for strongest baselines.

## Introduction and Theoretical Foundation

### Background

An LLM agent consists of a model wrapped in a *harness*: the prompts, tools, memory, and control flow that enable action. In **harness evolution**, an LLM optimizer reads failed trajectories of a frozen executor agent and edits the executor's harness over several rounds. However, the optimizer's own harness typically remains fixed—it is a hand-built pipeline or off-the-shelf coding agent.

### Key Question

> Can the optimizer evolve its own harness as well?

The authors find that simply allowing self-evolution *without* verification hurts performance: the optimizer writes untested lessons into its own harness, and no later round beats the initial harness on validation. However, when the optimizer can run a draft edit on a training task before submitting it, self-evolution achieves the best results.

### Theoretical Foundation

The paper formalizes harness evolution as **bilevel learning**:

- **Inner loop**: edits the executor's harness $h$ based on training failures
- **Outer loop**: updates the optimizer's own harness $g$ (prompts, skills, tools, hooks, notes)

The objectives are:

$$
\begin{array}{l l} 
\text{inner:} & h_{t+1}(g) \in \underset{h \in \mathcal{C}_t(g)}{\arg \min} \widehat{L}_{\text{train}}(h), \\ 
\text{outer:} & \underset{g}{\min} \mathbb{E}[\widehat{L}_{\text{val}}(h_{t+1}(g))]. 
\end{array}\tag{2}
$$

where $\widehat{L}_{\text{train}}$ and $\widehat{L}_{\text{val}}$ are empirical losses on training and validation sets, and $\mathcal{C}_t(g)$ is the set of candidate harnesses generated within budget $B_t$.

## Methodology

### VERSE Architecture

VERSE is added to an existing harness optimizer (the "host") and adds four components:

1. **Attribution**: Uses *trace minimization* to re-run subsets of failed trajectory steps, keeping only the few steps that still reproduce the same error. The median failed trajectory shrinks from 129 steps to 8.

2. **Verification tools** (three tools, each answering one question before submission):
   - **Verification tool**: runs a draft on up to three target tasks; a task counts as "fixed" only if it passes *twice* (since a single pass can be luck)
   - **Replay tool**: checks whether a recorded failure is reproducible by re-running the recorded trajectory
   - **Perturbation tool**: removes or replaces a suspected step in a reproducible trajectory to test whether that step causes the failure

3. **Training audit**: Tracks which failure modes persist across rounds, which tasks were fixed, which regressed, and which fixes were later undone.

4. **Optimizer self-evolution**: At the end of each round, the optimizer spends up to 30 turns reviewing results (validation accuracy, whether expected fixes were achieved, tool usage and errors) and edits its own harness $g$.

### Optimizer Harness Components

A self-evolving optimizer can rewrite five parts of its own harness:
- **Prompts**: system prompts
- **Skills**: instruction files added to the prompt every round
- **Tools**: functions it writes and can call
- **Hooks**: code that runs automatically at fixed points of the optimizer's loop
- **Notes**: memory of earlier rounds

Self-written code is loaded only after passing a safety check and test run. Everything else (evaluation protocol, budgets) stays fixed.

### Task Splits

To ensure generalization, the paper uses three disjoint task sets:
- **Training**: 110 tasks from SWE-rebench (Jan 2025–Feb 2026 releases)
- **Validation**: 50 tasks from the same pool
- **Test (in-distribution)**: 108 gradable tasks from March 2026 release (Python)
- **Test (OOD)**: 107 gradable tasks from July 2026 release (Go, Java, Python, Rust, TypeScript; only 20 in Python)

Validation selection mirrors model selection in ML: the checkpoint with best validation performance is kept, not the last one.

## Empirical Validation / Results

### Observation 1: Self-evolution without verification fails

**Table 1: Self-evolution with and without a verification tool** (SWE-rebench avg@3 accuracy, %, ±SEM):

| Optimizer | Self-evolving | Verified | $r^{\star}$ | Val-selected | Last |
|---|---|---|---|---|---|
| Blank (initial harness $h_0$) | - | - | - | 33.95±0.82 | 33.95±0.82 |
| Meta-Harness | ✗ | ✗ | 4 | 39.20±1.11 | 35.80±0.82 |
| Self-evolving (w/o verification) | ✓ | ✗ | 0 | 31.79±1.35 | 31.79±3.13 |
| Simple verification only | ✗ | ✓ | 5 | 39.81±1.41 | 37.65±2.74 |
| Self-evolving + simple verification | ✓ | ✓ | 6 | **41.05±0.82** | **41.05±0.82** |

Without verification, self-evolution made Meta-Harness *worse*: validation selected the initial harness ($r^\star = 0$). With verification, self-evolution reached the highest accuracy.

### Observation 2: What optimizers build for themselves

Across five executor models, the self-evolving optimizer evolved 51 artifacts (prompts, skills, tools, hooks, notes files), consistently building four kinds:
- **Attribution**: tools for finding and ranking causes of failures
- **Verification**: checks that drafts run correctly and improve the executor
- **Training audit**: notes tracking fixes, regressions, and undone fixes across rounds
- **Workflow**: rules and hooks changing how the optimizer works

### Main Results

**Table 2: Test accuracy of VERSE on SWE-rebench** (avg@3 accuracy, %, ±SEM):

| Method | $r^{\star}$ | In-dist Val-selected | In-dist Last | OOD Val-selected | OOD Last |
|---|---|---|---|---|---|
| Blank | - | 33.95±0.82 | 33.95±0.82 | 27.41±0.82 | 27.41±0.82 |
| Meta-Harness | 4 | 39.20±1.11 | 35.80±0.82 | 28.66±0.82 | 30.84±1.95 |
| + VERSE (w/ self-evolve) | 2 | **42.28±1.11** | 41.36±2.16 | **37.69±1.56** | 38.01±1.56 |
| AHE | 6 | 33.95±1.23 | 33.95±1.23 | 26.17±0.54 | 26.17±0.54 |
| + VERSE (w/ self-evolve) | 5 | 35.19±2.14 | 35.80±1.11 | 30.22±0.62 | 30.22±1.89 |
| Self-Harness | 6 | 29.63±0.53 | 29.63±0.53 | 29.28±2.72 | 29.28±2.72 |
| + VERSE (w/ self-evolve) | 1 | 34.88±1.35 | 33.33±1.93 | 30.22±0.82 | 31.78±1.95 |
| HarnessX | 0 | 32.72±1.11 | 34.26±2.45 | 27.41±1.25 | 32.09±1.65 |
| + VERSE (w/ self-evolve) | 3 | 35.80±0.82 | 36.11±1.60 | **37.69±2.43** | 32.71±0.93 |

Key findings:
- VERSE with self-evolution improves every baseline in all 16 reported comparisons (both test sets, both selected and final harness), by 0.6 to 10.3 points
- Best result: Meta-Harness host reaches 42.28% in-distribution (3.1 points above strongest baseline) and 37.69% OOD
- Without self-evolution, VERSE tools alone fall below baseline in 8 of 16 comparisons; with self-evolution, it stays above in all 16

### Ablation Results

**Table 3: Ablating VERSE on the AHE host** (in-distribution test set):

| Configuration | $r^*$ | Val-selected | Last |
|---|---|---|---|
| AHE + VERSE, full config | 2 | **37.96±2.14** | 32.10±0.31 |
| w/o verification during evolution | 3 | 37.65±1.35 | 35.49±0.31 |
| w/o attribution and minimization | 2 | 35.80±1.35 | 31.79±2.53 |
| w/o training audit only | 2 | 36.73±1.23 | 32.41±3.21 |
| AHE alone | 6 | 33.95±1.23 | 33.95±1.23 |

Removing attribution and minimization costs ~half of VERSE's gain (2.2 of 4.0 points); removing the training audit alone costs 1.2 points.

### What Verification Does During Evolution

- Only **1 in 10** tasks predicted to be fixed by submitted edits actually passed in the next training run
- Only **17 of 170** draft evaluations found that the draft fixed a targeted failing task
- Submitted edits fixed 148 validation tasks but **broke 134** (only counting changes that reproduced on rerun)
- **Half of all validation outcome changes did not reproduce** on rerun (hence the "pass twice" rule)
- Self-evolving optimizers added regression checks in 34 of 81 draft evaluations (vs. 2 of 89 for fixed optimizers)
- Self-evolving optimizers revised unsuccessful drafts in 79% of cases (vs. 57% without self-evolution)

## Theoretical Results

### Theorem 1 (Repair choice)

For a binary hypothesis test with two possible failure causes $W \in \{1, 2\}$:

$$
\mathbb{P}(\widehat{h} \neq h_W) = p_n^* + \varepsilon \leq \frac{1}{2}\beta_0 \prod_{j=1}^{n} \rho_j + \varepsilon.\tag{3}
$$

where $\beta_0$ is the Bhattacharyya coefficient of all evidence before checks, $\rho_j$ bounds the overlap of check $j$, and $\varepsilon$ is the optimizer's excess error over the best decision rule. If checks only re-analyze existing evidence (add no new information), then $p_n^* \geq \lambda(1-\lambda)\beta_0^2$ for every $n$.

**Implication**: Re-analyzing existing evidence cannot lower the minimum achievable error; verification adds *new execution evidence* that can separate the causes and tighten the bound.

### Theorem 2 (Finding and selecting a good harness)

$$
\mathbb{E}L_{\theta,D}(\widehat{h}) \leq \ell_0 + \underbrace{(1-\ell_0)\prod_{i=1}^{K}\left[1 - r_i(1-e_i)\right]}_{\text{no good harness found}} + \underbrace{\varepsilon_{\text{sel}}}_{\text{selection error}}.\tag{4}
$$

where $r_i$ is the probability of targeting the right failure, $e_i$ is the probability of submitting a wrong repair, and $K$ is the total number of proposals. The two factors multiply: better target selection helps little while repairs often fail—consistent with Observation 1.

## Theoretical and Practical Implications

### Theoretical Implications

- Provides the first formal framework for understanding *optimizer self-evolution* in harness evolution as bilevel learning
- Shows that execution-based verification is not merely helpful but theoretically necessary: reading trajectories alone leaves a minimum rate of wrong repairs when evidence is ambiguous
- The multiplication of $r_i$ and $(1-e_i)$ in Theorem 2 explains why self-evolution without verification fails: improving target selection alone cannot compensate for high repair-error rates

### Practical Implications

- **Harness evolution should be bilevel**: improving not only the agent's configuration but also the procedures that propose, test, and revise it
- **Execution checks are essential**: verification tools should be provided from the first round rather than left for the optimizer to discover
- **Validation-based selection** (keeping the best-validation harness, not the last one) is supported: it matched or beat the final-round harness in 15 of 18 runs
- **Regression checks matter**: edits can fix targets while breaking other tasks; self-evolving optimizers learn to add such checks
- The approach generalizes across programming languages (Python training → Go/Java/Rust/TypeScript OOD gains)

## Conclusion

VERSE demonstrates that harness evolution improves when the optimizer can also evolve its own procedures. The combination of execution-based verification (attribution, verification tools, training audit) with optimizer self-evolution yields consistent improvements over all four evaluated harness optimizers, both in-distribution and out-of-distribution.

**Key limitations and future directions**:
- The incremental benefit of self-evolution varies across hosts (it lowered selected accuracy in 5 of 6 comparisons vs. VERSE without self-evolution)
- Test accuracy is not available when choosing a host or round, so staying above baseline in every comparison matters in practice
- Future work could explore: more informative verification tools, better scheduling of verification calls, and understanding when self-evolution helps versus hurts

> **Broader principle**: Harness evolution should improve not only the agent's configuration, but also the procedures that propose, test, and revise it.

---

_Markdown view of https://picx.dev/p/Gv8uGC, served by PicX — AI-generated visual whiteboard summaries of research papers._
