Summary (Overview)

  • Core Research Question: When should an LLM agent be improved by evolving its harness (runtime code/prompts) versus training its weights? The paper develops a diagnostic rule based on failure composition analysis.
  • Key Finding: Harness evolution repairs process failures (blocked calls, loops, budget exhaustion) while content failures (poor delivered plans) require weight training. The gain from harness evolution can be internalized into weights when it changes what the model writes, but not when it changes what the model sees.
  • Quantitative Results: On DeepPlanning, self-evolving harness lifts Qwen3.5-4B held-out score from 0.16→0.30 and Qwen3.5-9B from 0.32→0.44. LoRA adapters trained on evolved-harness trajectories add +0.13 on held-out tasks for both sizes under the seed harness.
  • Key Distinction: On 4B, harness and weight levers stack; on 9B, they substitute (adapter alone matches full evolution line). The split depends on whether delivery is saturated or not.
  • Transfer Result: The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), but there the gain lives in what the model sees (two pages of history instead of one), so adapters trained on those trajectories do not add to it.

Introduction and Theoretical Foundation

Background and Motivation

A long-horizon LLM agent can be improved in two places:

  1. The harness: code and text around a frozen model (system prompts, memory tools, checkpoints, output gates, tools)
  2. The weights: fine-tuning on trajectories

Each lever has its own literature—self-evolving harnesses (Lin et al., 2026; Hu et al., 2024) and trajectory fine-tuning (Zeng et al., 2023; Chen et al., 2023)—but the question of which lever to pull when remains open.

The Central Hypothesis

The paper tests a diagnose-then-intervene rule: label failed trajectories by the first signal that fires, separating:

  • Process failures: blocked calls (F2), loops and exhausted step budgets (F3)
  • Content failures: a delivered plan that is simply poor (F4)

The hypothesis: Harness evolution repairs process failures, the behavior it instils can be trained into the weights, and content failures are what weight training is for.

Theoretical Foundations

Score identity: On DeepPlanning, an undelivered plan scores zero, so the composite score decomposes exactly as:

S=rqˉS = r \bar{q}

where rr is the delivery rate and qˉ\bar{q} is the mean score among delivered plans. A change from (r,qˉ)(r, \bar{q}) to (r′,qˉ′)(r', \bar{q}') decomposes as:

ΔS=qˉΔr+r′Δqˉ(2)\Delta S = \bar{q} \Delta r + r' \Delta \bar{q} \tag{2}

This identity is crucial: a lever that only moves rr is capped at qˉ\bar{q} (the quality of what the model already writes), while a lever that raises qˉ\bar{q} is worth r′Δqˉr' \Delta \bar{q}, more when delivery is already high.

Selection optimism bound (Proposition 1): When keeping the best of KK candidates with readings Xk=μk+εkX_k = \mu_k + \varepsilon_k where εk\varepsilon_k are zero-mean ss-sub-Gaussian:

E[Xk⋆−μk⋆]≤s2ln⁡K(1)\mathbb{E}\left[X_{k^\star} - \mu_{k^\star}\right] \leq s \sqrt{2 \ln K} \tag{1}

With single-rollout standard deviation σ=0.023\sigma = 0.023 and KK between 10–30 candidates, the bound is 0.05–0.06 (above the 0.035 resolution line), motivating the four-rollout protocol.


Methodology

1. Failure Taxonomy

Each trajectory is labeled post hoc by an automated checker. A trajectory scoring ≥ 0.5 is healthy; otherwise it receives the first signal that fires:

  • F1 (evidence): ≥ half of the plan's factual claims are unbacked, or the agent passed the exit gate with unresolved items
  • F2 (harness friction): a harness intervention blocked the agent (forced checkpoint unlock, or ≥ 10 blocked tool calls)
  • F3 (execution control): same tool call repeated ≥ 3 times, or ≥ 90% of step budget used, or harness termination
  • F4 (planning capacity): none of the above; a plan was delivered scoring < 0.5

2. Measurement Protocol

  • Benchmark: DeepPlanning (travel planning, deterministic programmatic scoring), 48 development tasks + 72 held-out tasks
  • Noise model: per-rollout standard deviation median 0.023 (worst 0.042); differences below 0.035 are "within noise"
  • Protocol: every configuration is a four-rollout mean; comparisons are same-night; fresh anchor follows every accepted change

3. Harness Self-Evolution Loop

A propose-verify loop in the style of AHE (Lin et al., 2026):

  1. Run current harness on 48 development tasks (4 rollouts)
  2. Build evidence package (score, healthy rate, F1–F4 counts, process metrics, worst cases) → hand to proposal model (Kimi K3)
  3. Proposer returns two candidate edits, each with manifest and rollback condition
  4. Evaluate each candidate (4 rollouts), judge against round's anchor by fixed house rules
  5. Retain better accepted candidate; measure fresh anchor

The editable surface is the harness definition only: prompt files, policy parameters, middleware, tools. The seed harness includes a state ledger, exit gate, mid-trajectory checkpoint, and four hard blocks.

4. Weight Training

  • Data: trajectories under evolved harness on 48 dev tasks (8–12 rollouts per task), keeping 3 best trajectories per task above threshold (0.5 for 4B, 0.6 for 9B)
  • LoRA: rank 32, α=64\alpha = 64, dropout 0.05, learning rate 10−410^{-4}, on attention projections (qkvo) or attention + MLP (qkvo+MLP), taken at end of first epoch
  • Full fine-tuning: 2×10−62 \times 10^{-6} with warm-up and cosine decay
  • Placebo control: adapter trained on answer-shuffled trajectories (assistant outputs permuted across samples)

Empirical Validation / Results

Harness Self-Evolution Works and Transfers

ModelDev GainHeld-out GainDelivery (before→after)Loops F3 (before→after)
Qwen3.5-4B+0.143+0.13359%→88%64%→37%
Qwen3.5-9B+0.056+0.11997% (already high)39%→17%

Key finding: F4 was never reduced by the loop. It grows from 10%→28% of 4B trajectories (as loops are converted) and stays at 25% for 9B, where it becomes the largest failure class.

Weight Training Internalizes the Harness Gain

Table 1 (key results, same-night evaluation):

QuantityWeightsQwen3.5-4B SeedQwen3.5-4B EvolvedQwen3.5-9B SeedQwen3.5-9B Evolved
Score (dev)Base0.1580.3010.3480.404
Score (dev)LoRA best0.3090.3510.3980.414
Score (held-out)Base0.1640.2970.3230.442
Score (held-out)LoRA best0.2900.3790.4540.449
Delivery (dev)Base52%90%96%98%
Delivery (dev)LoRA best78%95%95%100%

Four findings:

  1. The adapter transfers: on seed harness, better adapter lifts held-out by +0.126 (4B) and +0.131 (9B)
  2. Stacking on 4B, substitution on 9B: on evolved harness, 4B adapter adds +0.041 (dev) / +0.082 (held-out); 9B adapter adds +0.010 / +0.007 (within noise)
  3. Content, not adapter: placebo adapter (answer-shuffled) scores 0.158 below base on dev, 0.227 below on held-out
  4. Each adapter moves its own failure class: 4B adapter raises delivery (52%→78%) and cuts loops; 9B adapter cuts F4 from 23–28% to 5–7%

WebArena-Lite Transfer

  • 13 rounds, 10 accepted edits; dev anchor 0.234→0.375
  • One interface edit (two full pages of history instead of one) accounts for +0.12
  • Held-out: 0.256→0.344 (+0.088), delivery 49%→77%
  • Adapters did NOT transfer the gain: within noise of base under both harnesses, because the gain lives in what the model sees, not what it writes

Across Nine Evolution Lines

Seed measurements separate lines into three groups:

GroupSeed Plan QualityOutcome
5 lines (4B to hy3)≥ 0.27, F4 not dominantHarness paid +0.063 to +0.136
3 lines (MiniCPM5-2B, Spark-X2.5-4B, Ministral-3-8B)0.13–0.17Harness paid nothing
1 line (Gemma 4 12B)Delivery saturated, F4 dominantRule sends to weights

Theoretical and Practical Implications

Component Ablations: What the Harness Components Buy

  • Capability-granting components carry the gain: state ledger, distilled lessons file, invented tools
  • Enforcement mostly does not: exit gate, checkpoint, hard blocks are within noise (except 4B exit gate, which holds held-out delivery)

For 9B: removing the lessons file costs 0.049 (dev) / 0.076 (held-out); removing the ledger costs 0.088 / 0.145. For 4B: removing invented tools costs 0.071 / 0.090; removing lessons file costs 0.067 / 0.120.

Why the Loop Stops at F4

Four properties of the evidence package bound what the proposer can see:

  1. Package leads with intervention process metrics (components ablations show do not move score); F1–F4 counts come after; scorer's per-check failure composition never aggregated
  2. Proposal request is a single call with no tools
  3. Keys defining observed metrics live outside editable directories
  4. Accepted edits target what the package makes visible

The Diagnose-then-Intervene Rule

  1. Delivered plans score below ~0.18: harness lever pays nothing; don't train either
  2. Delivery well below saturation or F3 dominant: evolve the harness
  3. Delivery saturated and F4 dominant: train the weights
  4. F4 residual after both: expose scorer's failure composition to proposer
  5. Harness gain in hand: read what accepted edits changed—if they changed what the model does, adapters carry the gain; if from what the model sees, it stays in the harness

Conclusion

The paper establishes a principled, empirically-grounded rule for choosing between harness evolution and weight training for LLM agents:

  1. Measure the noise floor first, then delivery rate, quality among delivered plans, and F1–F4 composition
  2. Evolve the harness for process failures (F2, F3), which transfers to held-out tasks
  3. Train the weights for content failures (F4), which the harness loop cannot see or fix
  4. The split is size-dependent: on 4B the levers stack; on 9B they substitute

The key theoretical contribution is the failure-composition-based lever selection: gains from edits that change what the model writes can be trained into weights; gains from edits that change what the model sees stay in the harness.

Limitations

  • Taxonomy requires trajectory-level process signals and a task-level scorer
  • 4B ablation cells run with two rollouts (except exit-gate cell); placebo control exists for 9B only
  • No comparison with matched-budget test-time search baseline
  • Plan-to-JSON conversion is an LLM call (inflates scores ~0.025); two comparisons cross nights

Future Directions

  • Expose the scorer's failure composition to the proposer to push past the F4 ceiling
  • Train the three models where the harness lever paid nothing (MiniCPM5-2B, Spark-X2.5-4B, Ministral-3-8B) to test the rule's other branch
  • Extend the diagnose-then-intervene rule to broader benchmarks and more model families

Related papers