Summary (Overview)
- Core Research Question: When should an LLM agent be improved by evolving its harness (runtime code/prompts) versus training its weights? The paper develops a diagnostic rule based on failure composition analysis.
- Key Finding: Harness evolution repairs process failures (blocked calls, loops, budget exhaustion) while content failures (poor delivered plans) require weight training. The gain from harness evolution can be internalized into weights when it changes what the model writes, but not when it changes what the model sees.
- Quantitative Results: On DeepPlanning, self-evolving harness lifts Qwen3.5-4B held-out score from 0.16→0.30 and Qwen3.5-9B from 0.32→0.44. LoRA adapters trained on evolved-harness trajectories add +0.13 on held-out tasks for both sizes under the seed harness.
- Key Distinction: On 4B, harness and weight levers stack; on 9B, they substitute (adapter alone matches full evolution line). The split depends on whether delivery is saturated or not.
- Transfer Result: The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), but there the gain lives in what the model sees (two pages of history instead of one), so adapters trained on those trajectories do not add to it.
Introduction and Theoretical Foundation
Background and Motivation
A long-horizon LLM agent can be improved in two places:
- The harness: code and text around a frozen model (system prompts, memory tools, checkpoints, output gates, tools)
- The weights: fine-tuning on trajectories
Each lever has its own literature—self-evolving harnesses (Lin et al., 2026; Hu et al., 2024) and trajectory fine-tuning (Zeng et al., 2023; Chen et al., 2023)—but the question of which lever to pull when remains open.
The Central Hypothesis
The paper tests a diagnose-then-intervene rule: label failed trajectories by the first signal that fires, separating:
- Process failures: blocked calls (F2), loops and exhausted step budgets (F3)
- Content failures: a delivered plan that is simply poor (F4)
The hypothesis: Harness evolution repairs process failures, the behavior it instils can be trained into the weights, and content failures are what weight training is for.
Theoretical Foundations
Score identity: On DeepPlanning, an undelivered plan scores zero, so the composite score decomposes exactly as:
where is the delivery rate and is the mean score among delivered plans. A change from to decomposes as:
This identity is crucial: a lever that only moves is capped at (the quality of what the model already writes), while a lever that raises is worth , more when delivery is already high.
Selection optimism bound (Proposition 1): When keeping the best of candidates with readings where are zero-mean -sub-Gaussian:
With single-rollout standard deviation and between 10–30 candidates, the bound is 0.05–0.06 (above the 0.035 resolution line), motivating the four-rollout protocol.
Methodology
1. Failure Taxonomy
Each trajectory is labeled post hoc by an automated checker. A trajectory scoring ≥ 0.5 is healthy; otherwise it receives the first signal that fires:
- F1 (evidence): ≥ half of the plan's factual claims are unbacked, or the agent passed the exit gate with unresolved items
- F2 (harness friction): a harness intervention blocked the agent (forced checkpoint unlock, or ≥ 10 blocked tool calls)
- F3 (execution control): same tool call repeated ≥ 3 times, or ≥ 90% of step budget used, or harness termination
- F4 (planning capacity): none of the above; a plan was delivered scoring < 0.5
2. Measurement Protocol
- Benchmark: DeepPlanning (travel planning, deterministic programmatic scoring), 48 development tasks + 72 held-out tasks
- Noise model: per-rollout standard deviation median 0.023 (worst 0.042); differences below 0.035 are "within noise"
- Protocol: every configuration is a four-rollout mean; comparisons are same-night; fresh anchor follows every accepted change
3. Harness Self-Evolution Loop
A propose-verify loop in the style of AHE (Lin et al., 2026):
- Run current harness on 48 development tasks (4 rollouts)
- Build evidence package (score, healthy rate, F1–F4 counts, process metrics, worst cases) → hand to proposal model (Kimi K3)
- Proposer returns two candidate edits, each with manifest and rollback condition
- Evaluate each candidate (4 rollouts), judge against round's anchor by fixed house rules
- Retain better accepted candidate; measure fresh anchor
The editable surface is the harness definition only: prompt files, policy parameters, middleware, tools. The seed harness includes a state ledger, exit gate, mid-trajectory checkpoint, and four hard blocks.
4. Weight Training
- Data: trajectories under evolved harness on 48 dev tasks (8–12 rollouts per task), keeping 3 best trajectories per task above threshold (0.5 for 4B, 0.6 for 9B)
- LoRA: rank 32, , dropout 0.05, learning rate , on attention projections (qkvo) or attention + MLP (qkvo+MLP), taken at end of first epoch
- Full fine-tuning: with warm-up and cosine decay
- Placebo control: adapter trained on answer-shuffled trajectories (assistant outputs permuted across samples)
Empirical Validation / Results
Harness Self-Evolution Works and Transfers
| Model | Dev Gain | Held-out Gain | Delivery (before→after) | Loops F3 (before→after) |
|---|---|---|---|---|
| Qwen3.5-4B | +0.143 | +0.133 | 59%→88% | 64%→37% |
| Qwen3.5-9B | +0.056 | +0.119 | 97% (already high) | 39%→17% |
Key finding: F4 was never reduced by the loop. It grows from 10%→28% of 4B trajectories (as loops are converted) and stays at 25% for 9B, where it becomes the largest failure class.
Weight Training Internalizes the Harness Gain
Table 1 (key results, same-night evaluation):
| Quantity | Weights | Qwen3.5-4B Seed | Qwen3.5-4B Evolved | Qwen3.5-9B Seed | Qwen3.5-9B Evolved |
|---|---|---|---|---|---|
| Score (dev) | Base | 0.158 | 0.301 | 0.348 | 0.404 |
| Score (dev) | LoRA best | 0.309 | 0.351 | 0.398 | 0.414 |
| Score (held-out) | Base | 0.164 | 0.297 | 0.323 | 0.442 |
| Score (held-out) | LoRA best | 0.290 | 0.379 | 0.454 | 0.449 |
| Delivery (dev) | Base | 52% | 90% | 96% | 98% |
| Delivery (dev) | LoRA best | 78% | 95% | 95% | 100% |
Four findings:
- The adapter transfers: on seed harness, better adapter lifts held-out by +0.126 (4B) and +0.131 (9B)
- Stacking on 4B, substitution on 9B: on evolved harness, 4B adapter adds +0.041 (dev) / +0.082 (held-out); 9B adapter adds +0.010 / +0.007 (within noise)
- Content, not adapter: placebo adapter (answer-shuffled) scores 0.158 below base on dev, 0.227 below on held-out
- Each adapter moves its own failure class: 4B adapter raises delivery (52%→78%) and cuts loops; 9B adapter cuts F4 from 23–28% to 5–7%
WebArena-Lite Transfer
- 13 rounds, 10 accepted edits; dev anchor 0.234→0.375
- One interface edit (two full pages of history instead of one) accounts for +0.12
- Held-out: 0.256→0.344 (+0.088), delivery 49%→77%
- Adapters did NOT transfer the gain: within noise of base under both harnesses, because the gain lives in what the model sees, not what it writes
Across Nine Evolution Lines
Seed measurements separate lines into three groups:
| Group | Seed Plan Quality | Outcome |
|---|---|---|
| 5 lines (4B to hy3) | ≥ 0.27, F4 not dominant | Harness paid +0.063 to +0.136 |
| 3 lines (MiniCPM5-2B, Spark-X2.5-4B, Ministral-3-8B) | 0.13–0.17 | Harness paid nothing |
| 1 line (Gemma 4 12B) | Delivery saturated, F4 dominant | Rule sends to weights |
Theoretical and Practical Implications
Component Ablations: What the Harness Components Buy
- Capability-granting components carry the gain: state ledger, distilled lessons file, invented tools
- Enforcement mostly does not: exit gate, checkpoint, hard blocks are within noise (except 4B exit gate, which holds held-out delivery)
For 9B: removing the lessons file costs 0.049 (dev) / 0.076 (held-out); removing the ledger costs 0.088 / 0.145. For 4B: removing invented tools costs 0.071 / 0.090; removing lessons file costs 0.067 / 0.120.
Why the Loop Stops at F4
Four properties of the evidence package bound what the proposer can see:
- Package leads with intervention process metrics (components ablations show do not move score); F1–F4 counts come after; scorer's per-check failure composition never aggregated
- Proposal request is a single call with no tools
- Keys defining observed metrics live outside editable directories
- Accepted edits target what the package makes visible
The Diagnose-then-Intervene Rule
- Delivered plans score below ~0.18: harness lever pays nothing; don't train either
- Delivery well below saturation or F3 dominant: evolve the harness
- Delivery saturated and F4 dominant: train the weights
- F4 residual after both: expose scorer's failure composition to proposer
- Harness gain in hand: read what accepted edits changed—if they changed what the model does, adapters carry the gain; if from what the model sees, it stays in the harness
Conclusion
The paper establishes a principled, empirically-grounded rule for choosing between harness evolution and weight training for LLM agents:
- Measure the noise floor first, then delivery rate, quality among delivered plans, and F1–F4 composition
- Evolve the harness for process failures (F2, F3), which transfers to held-out tasks
- Train the weights for content failures (F4), which the harness loop cannot see or fix
- The split is size-dependent: on 4B the levers stack; on 9B they substitute
The key theoretical contribution is the failure-composition-based lever selection: gains from edits that change what the model writes can be trained into weights; gains from edits that change what the model sees stay in the harness.
Limitations
- Taxonomy requires trajectory-level process signals and a task-level scorer
- 4B ablation cells run with two rollouts (except exit-gate cell); placebo control exists for 9B only
- No comparison with matched-budget test-time search baseline
- Plan-to-JSON conversion is an LLM call (inflates scores ~0.025); two comparisons cross nights
Future Directions
- Expose the scorer's failure composition to the proposer to push past the F4 ceiling
- Train the three models where the harness lever paid nothing (MiniCPM5-2B, Spark-X2.5-4B, Ministral-3-8B) to test the rule's other branch
- Extend the diagnose-then-intervene rule to broader benchmarks and more model families
Related papers
- hacktrace: behavior-supervised detection of reward hacking during code generation
HACKTRACE detects reward hacking in coding agents from activations already computed during generation, cutting cheating from 85% to under 5% with only 8 ms overhead.
- Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
Tool calls are silently altered by execution-path hops in 12% of production shell invocations, and the IntAct repair protocol recovers 79.2% of resulting failures.
- One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents
Co-installed coding-agent skills that do the same job reduce the installed skill's usage by 19.9 percentage points without lowering task completion, a conflict decided at the first skill read.