Summary (Overview)
- VERSE (Verified Self-Evolving optimizer for agent harnesses) introduces a framework where an LLM optimizer can evolve both the executor's harness (prompts, tools, workflow) and its own optimizer harness, while keeping model weights fixed.
- Two key motivating observations: (1) optimizer self-evolution only improves performance when execution-based verification is available; (2) self-evolving optimizers naturally build tools for failure attribution, verification, training audits, and workflow control.
- VERSE provides four components: attribution (trace minimization), verification tools (verification, replay, perturbation), training audit, and optimizer self-evolution.
- The paper includes a theoretical framework viewing harness evolution as bilevel learning, proving that execution checks can lower the minimum rate of wrong repairs when evidence is ambiguous.
- Experiments on SWE-rebench show VERSE improves all four evaluated harness optimizers on held-out tasks and out-of-distribution tasks across five languages, achieving 42.3% (in-distribution) and 37.7% (OOD) accuracy vs. 39.2% and 29.3% for strongest baselines.
Introduction and Theoretical Foundation
Background
An LLM agent consists of a model wrapped in a harness: the prompts, tools, memory, and control flow that enable action. In harness evolution, an LLM optimizer reads failed trajectories of a frozen executor agent and edits the executor's harness over several rounds. However, the optimizer's own harness typically remains fixed—it is a hand-built pipeline or off-the-shelf coding agent.
Key Question
Can the optimizer evolve its own harness as well?
The authors find that simply allowing self-evolution without verification hurts performance: the optimizer writes untested lessons into its own harness, and no later round beats the initial harness on validation. However, when the optimizer can run a draft edit on a training task before submitting it, self-evolution achieves the best results.
Theoretical Foundation
The paper formalizes harness evolution as bilevel learning:
- Inner loop: edits the executor's harness based on training failures
- Outer loop: updates the optimizer's own harness (prompts, skills, tools, hooks, notes)
The objectives are:
where and are empirical losses on training and validation sets, and is the set of candidate harnesses generated within budget .
Methodology
VERSE Architecture
VERSE is added to an existing harness optimizer (the "host") and adds four components:
-
Attribution: Uses trace minimization to re-run subsets of failed trajectory steps, keeping only the few steps that still reproduce the same error. The median failed trajectory shrinks from 129 steps to 8.
-
Verification tools (three tools, each answering one question before submission):
- Verification tool: runs a draft on up to three target tasks; a task counts as "fixed" only if it passes twice (since a single pass can be luck)
- Replay tool: checks whether a recorded failure is reproducible by re-running the recorded trajectory
- Perturbation tool: removes or replaces a suspected step in a reproducible trajectory to test whether that step causes the failure
-
Training audit: Tracks which failure modes persist across rounds, which tasks were fixed, which regressed, and which fixes were later undone.
-
Optimizer self-evolution: At the end of each round, the optimizer spends up to 30 turns reviewing results (validation accuracy, whether expected fixes were achieved, tool usage and errors) and edits its own harness .
Optimizer Harness Components
A self-evolving optimizer can rewrite five parts of its own harness:
- Prompts: system prompts
- Skills: instruction files added to the prompt every round
- Tools: functions it writes and can call
- Hooks: code that runs automatically at fixed points of the optimizer's loop
- Notes: memory of earlier rounds
Self-written code is loaded only after passing a safety check and test run. Everything else (evaluation protocol, budgets) stays fixed.
Task Splits
To ensure generalization, the paper uses three disjoint task sets:
- Training: 110 tasks from SWE-rebench (Jan 2025–Feb 2026 releases)
- Validation: 50 tasks from the same pool
- Test (in-distribution): 108 gradable tasks from March 2026 release (Python)
- Test (OOD): 107 gradable tasks from July 2026 release (Go, Java, Python, Rust, TypeScript; only 20 in Python)
Validation selection mirrors model selection in ML: the checkpoint with best validation performance is kept, not the last one.
Empirical Validation / Results
Observation 1: Self-evolution without verification fails
Table 1: Self-evolution with and without a verification tool (SWE-rebench avg@3 accuracy, %, ±SEM):
| Optimizer | Self-evolving | Verified | Val-selected | Last | |
|---|---|---|---|---|---|
| Blank (initial harness ) | - | - | - | 33.95±0.82 | 33.95±0.82 |
| Meta-Harness | ✗ | ✗ | 4 | 39.20±1.11 | 35.80±0.82 |
| Self-evolving (w/o verification) | ✓ | ✗ | 0 | 31.79±1.35 | 31.79±3.13 |
| Simple verification only | ✗ | ✓ | 5 | 39.81±1.41 | 37.65±2.74 |
| Self-evolving + simple verification | ✓ | ✓ | 6 | 41.05±0.82 | 41.05±0.82 |
Without verification, self-evolution made Meta-Harness worse: validation selected the initial harness (). With verification, self-evolution reached the highest accuracy.
Observation 2: What optimizers build for themselves
Across five executor models, the self-evolving optimizer evolved 51 artifacts (prompts, skills, tools, hooks, notes files), consistently building four kinds:
- Attribution: tools for finding and ranking causes of failures
- Verification: checks that drafts run correctly and improve the executor
- Training audit: notes tracking fixes, regressions, and undone fixes across rounds
- Workflow: rules and hooks changing how the optimizer works
Main Results
Table 2: Test accuracy of VERSE on SWE-rebench (avg@3 accuracy, %, ±SEM):
| Method | In-dist Val-selected | In-dist Last | OOD Val-selected | OOD Last | |
|---|---|---|---|---|---|
| Blank | - | 33.95±0.82 | 33.95±0.82 | 27.41±0.82 | 27.41±0.82 |
| Meta-Harness | 4 | 39.20±1.11 | 35.80±0.82 | 28.66±0.82 | 30.84±1.95 |
| + VERSE (w/ self-evolve) | 2 | 42.28±1.11 | 41.36±2.16 | 37.69±1.56 | 38.01±1.56 |
| AHE | 6 | 33.95±1.23 | 33.95±1.23 | 26.17±0.54 | 26.17±0.54 |
| + VERSE (w/ self-evolve) | 5 | 35.19±2.14 | 35.80±1.11 | 30.22±0.62 | 30.22±1.89 |
| Self-Harness | 6 | 29.63±0.53 | 29.63±0.53 | 29.28±2.72 | 29.28±2.72 |
| + VERSE (w/ self-evolve) | 1 | 34.88±1.35 | 33.33±1.93 | 30.22±0.82 | 31.78±1.95 |
| HarnessX | 0 | 32.72±1.11 | 34.26±2.45 | 27.41±1.25 | 32.09±1.65 |
| + VERSE (w/ self-evolve) | 3 | 35.80±0.82 | 36.11±1.60 | 37.69±2.43 | 32.71±0.93 |
Key findings:
- VERSE with self-evolution improves every baseline in all 16 reported comparisons (both test sets, both selected and final harness), by 0.6 to 10.3 points
- Best result: Meta-Harness host reaches 42.28% in-distribution (3.1 points above strongest baseline) and 37.69% OOD
- Without self-evolution, VERSE tools alone fall below baseline in 8 of 16 comparisons; with self-evolution, it stays above in all 16
Ablation Results
Table 3: Ablating VERSE on the AHE host (in-distribution test set):
| Configuration | Val-selected | Last | |
|---|---|---|---|
| AHE + VERSE, full config | 2 | 37.96±2.14 | 32.10±0.31 |
| w/o verification during evolution | 3 | 37.65±1.35 | 35.49±0.31 |
| w/o attribution and minimization | 2 | 35.80±1.35 | 31.79±2.53 |
| w/o training audit only | 2 | 36.73±1.23 | 32.41±3.21 |
| AHE alone | 6 | 33.95±1.23 | 33.95±1.23 |
Removing attribution and minimization costs ~half of VERSE's gain (2.2 of 4.0 points); removing the training audit alone costs 1.2 points.
What Verification Does During Evolution
- Only 1 in 10 tasks predicted to be fixed by submitted edits actually passed in the next training run
- Only 17 of 170 draft evaluations found that the draft fixed a targeted failing task
- Submitted edits fixed 148 validation tasks but broke 134 (only counting changes that reproduced on rerun)
- Half of all validation outcome changes did not reproduce on rerun (hence the "pass twice" rule)
- Self-evolving optimizers added regression checks in 34 of 81 draft evaluations (vs. 2 of 89 for fixed optimizers)
- Self-evolving optimizers revised unsuccessful drafts in 79% of cases (vs. 57% without self-evolution)
Theoretical Results
Theorem 1 (Repair choice)
For a binary hypothesis test with two possible failure causes :
where is the Bhattacharyya coefficient of all evidence before checks, bounds the overlap of check , and is the optimizer's excess error over the best decision rule. If checks only re-analyze existing evidence (add no new information), then for every .
Implication: Re-analyzing existing evidence cannot lower the minimum achievable error; verification adds new execution evidence that can separate the causes and tighten the bound.
Theorem 2 (Finding and selecting a good harness)
where is the probability of targeting the right failure, is the probability of submitting a wrong repair, and is the total number of proposals. The two factors multiply: better target selection helps little while repairs often fail—consistent with Observation 1.
Theoretical and Practical Implications
Theoretical Implications
- Provides the first formal framework for understanding optimizer self-evolution in harness evolution as bilevel learning
- Shows that execution-based verification is not merely helpful but theoretically necessary: reading trajectories alone leaves a minimum rate of wrong repairs when evidence is ambiguous
- The multiplication of and in Theorem 2 explains why self-evolution without verification fails: improving target selection alone cannot compensate for high repair-error rates
Practical Implications
- Harness evolution should be bilevel: improving not only the agent's configuration but also the procedures that propose, test, and revise it
- Execution checks are essential: verification tools should be provided from the first round rather than left for the optimizer to discover
- Validation-based selection (keeping the best-validation harness, not the last one) is supported: it matched or beat the final-round harness in 15 of 18 runs
- Regression checks matter: edits can fix targets while breaking other tasks; self-evolving optimizers learn to add such checks
- The approach generalizes across programming languages (Python training → Go/Java/Rust/TypeScript OOD gains)
Conclusion
VERSE demonstrates that harness evolution improves when the optimizer can also evolve its own procedures. The combination of execution-based verification (attribution, verification tools, training audit) with optimizer self-evolution yields consistent improvements over all four evaluated harness optimizers, both in-distribution and out-of-distribution.
Key limitations and future directions:
- The incremental benefit of self-evolution varies across hosts (it lowered selected accuracy in 5 of 6 comparisons vs. VERSE without self-evolution)
- Test accuracy is not available when choosing a host or round, so staying above baseline in every comparison matters in practice
- Future work could explore: more informative verification tools, better scheduling of verification calls, and understanding when self-evolution helps versus hurts
Broader principle: Harness evolution should improve not only the agent's configuration, but also the procedures that propose, test, and revise it.
Related papers
- Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild
Agent harnesses remain substantially undertested, with less than half of LLM-dependent code covered, and HarnessTester's contract-aware test generation boosts coverage and mutation scores by up to 95% while finding 88 previously-unknown bugs.
- Agent Plasticity: Measuring Self-Improvement Through Experience
Introducing agent plasticity as a metric for learning efficiency reveals frontier models differ sharply in converting experience into persistent, generalizable performance gains.
- TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
TESTJACK refutes 34.4% of benchmark-passing LLM coding trials by evolving tests from prompt requirements, cutting reported resolution rates from 50.6% to 33.2%.