# Harness Evolution Hits a Ceiling: When Weight Training Should Begin

> Harness evolution fixes process failures like loops and blocked calls, while weight training fixes content failures, with gains transferring only when edits change what the model writes.

- **Source:** [arXiv](https://arxiv.org/abs/2610.11655)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/p5Hz36
- **Whiteboard:** https://picx.dev/p/p5Hz36/image

## Summary

## Summary (Overview)

- **Core Research Question**: When should an LLM agent be improved by evolving its harness (runtime code/prompts) versus training its weights? The paper develops a diagnostic rule based on failure composition analysis.
- **Key Finding**: Harness evolution repairs *process failures* (blocked calls, loops, budget exhaustion) while *content failures* (poor delivered plans) require weight training. The gain from harness evolution can be internalized into weights when it changes *what the model writes*, but not when it changes *what the model sees*.
- **Quantitative Results**: On DeepPlanning, self-evolving harness lifts Qwen3.5-4B held-out score from 0.16→0.30 and Qwen3.5-9B from 0.32→0.44. LoRA adapters trained on evolved-harness trajectories add +0.13 on held-out tasks for both sizes under the seed harness.
- **Key Distinction**: On 4B, harness and weight levers *stack*; on 9B, they *substitute* (adapter alone matches full evolution line). The split depends on whether delivery is saturated or not.
- **Transfer Result**: The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), but there the gain lives in *what the model sees* (two pages of history instead of one), so adapters trained on those trajectories do not add to it.

---

## Introduction and Theoretical Foundation

### Background and Motivation

A long-horizon LLM agent can be improved in two places:
1. **The harness**: code and text around a frozen model (system prompts, memory tools, checkpoints, output gates, tools)
2. **The weights**: fine-tuning on trajectories

Each lever has its own literature—self-evolving harnesses (Lin et al., 2026; Hu et al., 2024) and trajectory fine-tuning (Zeng et al., 2023; Chen et al., 2023)—but the question of *which lever to pull when* remains open.

### The Central Hypothesis

The paper tests a **diagnose-then-intervene rule**: label failed trajectories by the *first signal that fires*, separating:

- **Process failures**: blocked calls (F2), loops and exhausted step budgets (F3)
- **Content failures**: a delivered plan that is simply poor (F4)

The hypothesis: *Harness evolution repairs process failures, the behavior it instils can be trained into the weights, and content failures are what weight training is for.*

### Theoretical Foundations

**Score identity**: On DeepPlanning, an undelivered plan scores zero, so the composite score decomposes exactly as:

$$S = r \bar{q}$$

where $r$ is the delivery rate and $\bar{q}$ is the mean score among delivered plans. A change from $(r, \bar{q})$ to $(r', \bar{q}')$ decomposes as:

$$\Delta S = \bar{q} \Delta r + r' \Delta \bar{q} \tag{2}$$

This identity is crucial: a lever that only moves $r$ is capped at $\bar{q}$ (the quality of what the model already writes), while a lever that raises $\bar{q}$ is worth $r' \Delta \bar{q}$, more when delivery is already high.

**Selection optimism bound** (Proposition 1): When keeping the best of $K$ candidates with readings $X_k = \mu_k + \varepsilon_k$ where $\varepsilon_k$ are zero-mean $s$-sub-Gaussian:

$$\mathbb{E}\left[X_{k^\star} - \mu_{k^\star}\right] \leq s \sqrt{2 \ln K} \tag{1}$$

With single-rollout standard deviation $\sigma = 0.023$ and $K$ between 10–30 candidates, the bound is 0.05–0.06 (above the 0.035 resolution line), motivating the four-rollout protocol.

---

## Methodology

### 1. Failure Taxonomy

Each trajectory is labeled post hoc by an automated checker. A trajectory scoring ≥ 0.5 is healthy; otherwise it receives the **first** signal that fires:

- **F1 (evidence)**: ≥ half of the plan's factual claims are unbacked, or the agent passed the exit gate with unresolved items
- **F2 (harness friction)**: a harness intervention blocked the agent (forced checkpoint unlock, or ≥ 10 blocked tool calls)
- **F3 (execution control)**: same tool call repeated ≥ 3 times, or ≥ 90% of step budget used, or harness termination
- **F4 (planning capacity)**: none of the above; a plan was delivered scoring < 0.5

### 2. Measurement Protocol

- **Benchmark**: DeepPlanning (travel planning, deterministic programmatic scoring), 48 development tasks + 72 held-out tasks
- **Noise model**: per-rollout standard deviation median 0.023 (worst 0.042); differences below 0.035 are "within noise"
- **Protocol**: every configuration is a four-rollout mean; comparisons are same-night; fresh anchor follows every accepted change

### 3. Harness Self-Evolution Loop

A propose-verify loop in the style of AHE (Lin et al., 2026):

1. Run current harness on 48 development tasks (4 rollouts)
2. Build evidence package (score, healthy rate, F1–F4 counts, process metrics, worst cases) → hand to proposal model (Kimi K3)
3. Proposer returns two candidate edits, each with manifest and rollback condition
4. Evaluate each candidate (4 rollouts), judge against round's anchor by fixed house rules
5. Retain better accepted candidate; measure fresh anchor

The editable surface is the harness definition only: prompt files, policy parameters, middleware, tools. The seed harness includes a state ledger, exit gate, mid-trajectory checkpoint, and four hard blocks.

### 4. Weight Training

- **Data**: trajectories under evolved harness on 48 dev tasks (8–12 rollouts per task), keeping 3 best trajectories per task above threshold (0.5 for 4B, 0.6 for 9B)
- **LoRA**: rank 32, $\alpha = 64$, dropout 0.05, learning rate $10^{-4}$, on attention projections (qkvo) or attention + MLP (qkvo+MLP), taken at end of first epoch
- **Full fine-tuning**: $2 \times 10^{-6}$ with warm-up and cosine decay
- **Placebo control**: adapter trained on answer-shuffled trajectories (assistant outputs permuted across samples)

---

## Empirical Validation / Results

### Harness Self-Evolution Works and Transfers

| Model | Dev Gain | Held-out Gain | Delivery (before→after) | Loops F3 (before→after) |
|-------|----------|---------------|------------------------|------------------------|
| Qwen3.5-4B | +0.143 | +0.133 | 59%→88% | 64%→37% |
| Qwen3.5-9B | +0.056 | +0.119 | 97% (already high) | 39%→17% |

Key finding: **F4 was never reduced** by the loop. It grows from 10%→28% of 4B trajectories (as loops are converted) and stays at 25% for 9B, where it becomes the largest failure class.

### Weight Training Internalizes the Harness Gain

**Table 1 (key results, same-night evaluation):**

| Quantity | Weights | Qwen3.5-4B Seed | Qwen3.5-4B Evolved | Qwen3.5-9B Seed | Qwen3.5-9B Evolved |
|----------|---------|-----------------|--------------------|-----------------|--------------------|
| Score (dev) | Base | 0.158 | 0.301 | 0.348 | 0.404 |
| Score (dev) | LoRA best | 0.309 | 0.351 | 0.398 | 0.414 |
| Score (held-out) | Base | 0.164 | 0.297 | 0.323 | 0.442 |
| Score (held-out) | LoRA best | 0.290 | 0.379 | 0.454 | 0.449 |
| Delivery (dev) | Base | 52% | 90% | 96% | 98% |
| Delivery (dev) | LoRA best | 78% | 95% | 95% | 100% |

**Four findings:**
1. **The adapter transfers**: on seed harness, better adapter lifts held-out by +0.126 (4B) and +0.131 (9B)
2. **Stacking on 4B, substitution on 9B**: on evolved harness, 4B adapter adds +0.041 (dev) / +0.082 (held-out); 9B adapter adds +0.010 / +0.007 (within noise)
3. **Content, not adapter**: placebo adapter (answer-shuffled) scores 0.158 below base on dev, 0.227 below on held-out
4. **Each adapter moves its own failure class**: 4B adapter raises delivery (52%→78%) and cuts loops; 9B adapter cuts F4 from 23–28% to 5–7%

### WebArena-Lite Transfer

- 13 rounds, 10 accepted edits; dev anchor 0.234→0.375
- One interface edit (two full pages of history instead of one) accounts for +0.12
- Held-out: 0.256→0.344 (+0.088), delivery 49%→77%
- **Adapters did NOT transfer the gain**: within noise of base under both harnesses, because the gain lives in *what the model sees*, not what it writes

### Across Nine Evolution Lines

Seed measurements separate lines into three groups:

| Group | Seed Plan Quality | Outcome |
|-------|-------------------|---------|
| 5 lines (4B to hy3) | ≥ 0.27, F4 not dominant | Harness paid +0.063 to +0.136 |
| 3 lines (MiniCPM5-2B, Spark-X2.5-4B, Ministral-3-8B) | 0.13–0.17 | Harness paid nothing |
| 1 line (Gemma 4 12B) | Delivery saturated, F4 dominant | Rule sends to weights |

---

## Theoretical and Practical Implications

### Component Ablations: What the Harness Components Buy

- **Capability-granting components carry the gain**: state ledger, distilled lessons file, invented tools
- **Enforcement mostly does not**: exit gate, checkpoint, hard blocks are within noise (except 4B exit gate, which holds held-out delivery)

For 9B: removing the lessons file costs 0.049 (dev) / 0.076 (held-out); removing the ledger costs 0.088 / 0.145. For 4B: removing invented tools costs 0.071 / 0.090; removing lessons file costs 0.067 / 0.120.

### Why the Loop Stops at F4

Four properties of the evidence package bound what the proposer can see:
1. Package leads with intervention process metrics (components ablations show do not move score); F1–F4 counts come after; scorer's per-check failure composition never aggregated
2. Proposal request is a single call with no tools
3. Keys defining observed metrics live outside editable directories
4. Accepted edits target what the package makes visible

### The Diagnose-then-Intervene Rule

1. **Delivered plans score below ~0.18**: harness lever pays nothing; don't train either
2. **Delivery well below saturation or F3 dominant**: evolve the harness
3. **Delivery saturated and F4 dominant**: train the weights
4. **F4 residual after both**: expose scorer's failure composition to proposer
5. **Harness gain in hand**: read what accepted edits changed—if they changed *what the model does*, adapters carry the gain; if from *what the model sees*, it stays in the harness

---

## Conclusion

The paper establishes a principled, empirically-grounded rule for choosing between harness evolution and weight training for LLM agents:

1. **Measure the noise floor first**, then delivery rate, quality among delivered plans, and F1–F4 composition
2. **Evolve the harness for process failures** (F2, F3), which transfers to held-out tasks
3. **Train the weights for content failures** (F4), which the harness loop cannot see or fix
4. **The split is size-dependent**: on 4B the levers stack; on 9B they substitute

The key theoretical contribution is the **failure-composition-based lever selection**: gains from edits that change *what the model writes* can be trained into weights; gains from edits that change *what the model sees* stay in the harness.

### Limitations

- Taxonomy requires trajectory-level process signals and a task-level scorer
- 4B ablation cells run with two rollouts (except exit-gate cell); placebo control exists for 9B only
- No comparison with matched-budget test-time search baseline
- Plan-to-JSON conversion is an LLM call (inflates scores ~0.025); two comparisons cross nights

### Future Directions

- Expose the scorer's failure composition to the proposer to push past the F4 ceiling
- Train the three models where the harness lever paid nothing (MiniCPM5-2B, Spark-X2.5-4B, Ministral-3-8B) to test the rule's other branch
- Extend the diagnose-then-intervene rule to broader benchmarks and more model families

---

_Markdown view of https://picx.dev/p/p5Hz36, served by PicX — AI-generated visual whiteboard summaries of research papers._
