Summary (Overview)

  • VERSE (Verified Self-Evolving optimizer for agent harnesses) introduces a framework where an LLM optimizer can evolve both the executor's harness (prompts, tools, workflow) and its own optimizer harness, while keeping model weights fixed.
  • Two key motivating observations: (1) optimizer self-evolution only improves performance when execution-based verification is available; (2) self-evolving optimizers naturally build tools for failure attribution, verification, training audits, and workflow control.
  • VERSE provides four components: attribution (trace minimization), verification tools (verification, replay, perturbation), training audit, and optimizer self-evolution.
  • The paper includes a theoretical framework viewing harness evolution as bilevel learning, proving that execution checks can lower the minimum rate of wrong repairs when evidence is ambiguous.
  • Experiments on SWE-rebench show VERSE improves all four evaluated harness optimizers on held-out tasks and out-of-distribution tasks across five languages, achieving 42.3% (in-distribution) and 37.7% (OOD) accuracy vs. 39.2% and 29.3% for strongest baselines.

Introduction and Theoretical Foundation

Background

An LLM agent consists of a model wrapped in a harness: the prompts, tools, memory, and control flow that enable action. In harness evolution, an LLM optimizer reads failed trajectories of a frozen executor agent and edits the executor's harness over several rounds. However, the optimizer's own harness typically remains fixed—it is a hand-built pipeline or off-the-shelf coding agent.

Key Question

Can the optimizer evolve its own harness as well?

The authors find that simply allowing self-evolution without verification hurts performance: the optimizer writes untested lessons into its own harness, and no later round beats the initial harness on validation. However, when the optimizer can run a draft edit on a training task before submitting it, self-evolution achieves the best results.

Theoretical Foundation

The paper formalizes harness evolution as bilevel learning:

  • Inner loop: edits the executor's harness hh based on training failures
  • Outer loop: updates the optimizer's own harness gg (prompts, skills, tools, hooks, notes)

The objectives are:

inner:ht+1(g)∈arg⁡min⁡h∈Ct(g)L^train(h),outer:min⁡gE[L^val(ht+1(g))].(2)\begin{array}{l l} \text{inner:} & h_{t+1}(g) \in \underset{h \in \mathcal{C}_t(g)}{\arg \min} \widehat{L}_{\text{train}}(h), \\ \text{outer:} & \underset{g}{\min} \mathbb{E}[\widehat{L}_{\text{val}}(h_{t+1}(g))]. \end{array}\tag{2}

where L^train\widehat{L}_{\text{train}} and L^val\widehat{L}_{\text{val}} are empirical losses on training and validation sets, and Ct(g)\mathcal{C}_t(g) is the set of candidate harnesses generated within budget BtB_t.

Methodology

VERSE Architecture

VERSE is added to an existing harness optimizer (the "host") and adds four components:

  1. Attribution: Uses trace minimization to re-run subsets of failed trajectory steps, keeping only the few steps that still reproduce the same error. The median failed trajectory shrinks from 129 steps to 8.

  2. Verification tools (three tools, each answering one question before submission):

    • Verification tool: runs a draft on up to three target tasks; a task counts as "fixed" only if it passes twice (since a single pass can be luck)
    • Replay tool: checks whether a recorded failure is reproducible by re-running the recorded trajectory
    • Perturbation tool: removes or replaces a suspected step in a reproducible trajectory to test whether that step causes the failure
  3. Training audit: Tracks which failure modes persist across rounds, which tasks were fixed, which regressed, and which fixes were later undone.

  4. Optimizer self-evolution: At the end of each round, the optimizer spends up to 30 turns reviewing results (validation accuracy, whether expected fixes were achieved, tool usage and errors) and edits its own harness gg.

Optimizer Harness Components

A self-evolving optimizer can rewrite five parts of its own harness:

  • Prompts: system prompts
  • Skills: instruction files added to the prompt every round
  • Tools: functions it writes and can call
  • Hooks: code that runs automatically at fixed points of the optimizer's loop
  • Notes: memory of earlier rounds

Self-written code is loaded only after passing a safety check and test run. Everything else (evaluation protocol, budgets) stays fixed.

Task Splits

To ensure generalization, the paper uses three disjoint task sets:

  • Training: 110 tasks from SWE-rebench (Jan 2025–Feb 2026 releases)
  • Validation: 50 tasks from the same pool
  • Test (in-distribution): 108 gradable tasks from March 2026 release (Python)
  • Test (OOD): 107 gradable tasks from July 2026 release (Go, Java, Python, Rust, TypeScript; only 20 in Python)

Validation selection mirrors model selection in ML: the checkpoint with best validation performance is kept, not the last one.

Empirical Validation / Results

Observation 1: Self-evolution without verification fails

Table 1: Self-evolution with and without a verification tool (SWE-rebench avg@3 accuracy, %, ±SEM):

OptimizerSelf-evolvingVerifiedr⋆r^{\star}Val-selectedLast
Blank (initial harness h0h_0)---33.95±0.8233.95±0.82
Meta-Harness✗✗439.20±1.1135.80±0.82
Self-evolving (w/o verification)✓✗031.79±1.3531.79±3.13
Simple verification only✗✓539.81±1.4137.65±2.74
Self-evolving + simple verification✓✓641.05±0.8241.05±0.82

Without verification, self-evolution made Meta-Harness worse: validation selected the initial harness (r⋆=0r^\star = 0). With verification, self-evolution reached the highest accuracy.

Observation 2: What optimizers build for themselves

Across five executor models, the self-evolving optimizer evolved 51 artifacts (prompts, skills, tools, hooks, notes files), consistently building four kinds:

  • Attribution: tools for finding and ranking causes of failures
  • Verification: checks that drafts run correctly and improve the executor
  • Training audit: notes tracking fixes, regressions, and undone fixes across rounds
  • Workflow: rules and hooks changing how the optimizer works

Main Results

Table 2: Test accuracy of VERSE on SWE-rebench (avg@3 accuracy, %, ±SEM):

Methodr⋆r^{\star}In-dist Val-selectedIn-dist LastOOD Val-selectedOOD Last
Blank-33.95±0.8233.95±0.8227.41±0.8227.41±0.82
Meta-Harness439.20±1.1135.80±0.8228.66±0.8230.84±1.95
+ VERSE (w/ self-evolve)242.28±1.1141.36±2.1637.69±1.5638.01±1.56
AHE633.95±1.2333.95±1.2326.17±0.5426.17±0.54
+ VERSE (w/ self-evolve)535.19±2.1435.80±1.1130.22±0.6230.22±1.89
Self-Harness629.63±0.5329.63±0.5329.28±2.7229.28±2.72
+ VERSE (w/ self-evolve)134.88±1.3533.33±1.9330.22±0.8231.78±1.95
HarnessX032.72±1.1134.26±2.4527.41±1.2532.09±1.65
+ VERSE (w/ self-evolve)335.80±0.8236.11±1.6037.69±2.4332.71±0.93

Key findings:

  • VERSE with self-evolution improves every baseline in all 16 reported comparisons (both test sets, both selected and final harness), by 0.6 to 10.3 points
  • Best result: Meta-Harness host reaches 42.28% in-distribution (3.1 points above strongest baseline) and 37.69% OOD
  • Without self-evolution, VERSE tools alone fall below baseline in 8 of 16 comparisons; with self-evolution, it stays above in all 16

Ablation Results

Table 3: Ablating VERSE on the AHE host (in-distribution test set):

Configurationr∗r^*Val-selectedLast
AHE + VERSE, full config237.96±2.1432.10±0.31
w/o verification during evolution337.65±1.3535.49±0.31
w/o attribution and minimization235.80±1.3531.79±2.53
w/o training audit only236.73±1.2332.41±3.21
AHE alone633.95±1.2333.95±1.23

Removing attribution and minimization costs ~half of VERSE's gain (2.2 of 4.0 points); removing the training audit alone costs 1.2 points.

What Verification Does During Evolution

  • Only 1 in 10 tasks predicted to be fixed by submitted edits actually passed in the next training run
  • Only 17 of 170 draft evaluations found that the draft fixed a targeted failing task
  • Submitted edits fixed 148 validation tasks but broke 134 (only counting changes that reproduced on rerun)
  • Half of all validation outcome changes did not reproduce on rerun (hence the "pass twice" rule)
  • Self-evolving optimizers added regression checks in 34 of 81 draft evaluations (vs. 2 of 89 for fixed optimizers)
  • Self-evolving optimizers revised unsuccessful drafts in 79% of cases (vs. 57% without self-evolution)

Theoretical Results

Theorem 1 (Repair choice)

For a binary hypothesis test with two possible failure causes W∈{1,2}W \in \{1, 2\}:

P(h^≠hW)=pn∗+ε≤12β0∏j=1nρj+ε.(3)\mathbb{P}(\widehat{h} \neq h_W) = p_n^* + \varepsilon \leq \frac{1}{2}\beta_0 \prod_{j=1}^{n} \rho_j + \varepsilon.\tag{3}

where β0\beta_0 is the Bhattacharyya coefficient of all evidence before checks, ρj\rho_j bounds the overlap of check jj, and ε\varepsilon is the optimizer's excess error over the best decision rule. If checks only re-analyze existing evidence (add no new information), then pn∗≥λ(1−λ)β02p_n^* \geq \lambda(1-\lambda)\beta_0^2 for every nn.

Implication: Re-analyzing existing evidence cannot lower the minimum achievable error; verification adds new execution evidence that can separate the causes and tighten the bound.

Theorem 2 (Finding and selecting a good harness)

ELθ,D(h^)≤ℓ0+(1−ℓ0)∏i=1K[1−ri(1−ei)]⏟no good harness found+εsel⏟selection error.(4)\mathbb{E}L_{\theta,D}(\widehat{h}) \leq \ell_0 + \underbrace{(1-\ell_0)\prod_{i=1}^{K}\left[1 - r_i(1-e_i)\right]}_{\text{no good harness found}} + \underbrace{\varepsilon_{\text{sel}}}_{\text{selection error}}.\tag{4}

where rir_i is the probability of targeting the right failure, eie_i is the probability of submitting a wrong repair, and KK is the total number of proposals. The two factors multiply: better target selection helps little while repairs often fail—consistent with Observation 1.

Theoretical and Practical Implications

Theoretical Implications

  • Provides the first formal framework for understanding optimizer self-evolution in harness evolution as bilevel learning
  • Shows that execution-based verification is not merely helpful but theoretically necessary: reading trajectories alone leaves a minimum rate of wrong repairs when evidence is ambiguous
  • The multiplication of rir_i and (1−ei)(1-e_i) in Theorem 2 explains why self-evolution without verification fails: improving target selection alone cannot compensate for high repair-error rates

Practical Implications

  • Harness evolution should be bilevel: improving not only the agent's configuration but also the procedures that propose, test, and revise it
  • Execution checks are essential: verification tools should be provided from the first round rather than left for the optimizer to discover
  • Validation-based selection (keeping the best-validation harness, not the last one) is supported: it matched or beat the final-round harness in 15 of 18 runs
  • Regression checks matter: edits can fix targets while breaking other tasks; self-evolving optimizers learn to add such checks
  • The approach generalizes across programming languages (Python training → Go/Java/Rust/TypeScript OOD gains)

Conclusion

VERSE demonstrates that harness evolution improves when the optimizer can also evolve its own procedures. The combination of execution-based verification (attribution, verification tools, training audit) with optimizer self-evolution yields consistent improvements over all four evaluated harness optimizers, both in-distribution and out-of-distribution.

Key limitations and future directions:

  • The incremental benefit of self-evolution varies across hosts (it lowered selected accuracy in 5 of 6 comparisons vs. VERSE without self-evolution)
  • Test accuracy is not available when choosing a host or round, so staying above baseline in every comparison matters in practice
  • Future work could explore: more informative verification tools, better scheduling of verification calls, and understanding when self-evolution helps versus hurts

Broader principle: Harness evolution should improve not only the agent's configuration, but also the procedures that propose, test, and revise it.

Related papers