# EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?

> EVOHARNESSBENCH shows that expanding an agent's tool, skill, or agent harness alone can degrade performance on previously solved tasks, inducing forgetting up to 34.7%.

- **Source:** [arXiv](https://arxiv.org/abs/2609.04280)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/L2X26Z
- **Whiteboard:** https://picx.dev/p/L2X26Z/image

## Summary

## Summary (Overview)

- **EVOHARNESSBENCH** is a new benchmark for evaluating LLM-based agents under *harness evolution*—the controlled, cumulative growth of externally supplied tools, skills, and specialist agents.
- The benchmark contains **17 multi-stage harness streams** (3–6 stages each) with **802 unique tasks**, **520 tools**, **42 latent reference skills**, and **62 specialist agents**, constructed deterministically from verifier-based benchmarks (EnterpriseOps-Gym and Agents' Last Exam).
- Two evaluation settings are introduced: **deployment evaluation** (no persistent state; isolates harness-induced forgetting) and **self-evolving adaptation evaluation** (persistent artifacts carried across stages).
- Key findings reveal: (1) harness expansion alone can degrade performance on previously solved tasks (**harness-induced forgetting**, up to −46.4% for agents); (2) self-evolving adaptation gains are inconsistent across stages, axes, and environments; (3) retention and adaptation are often in tension.
- The work establishes harness evolution as a distinct challenge, motivating **bi-level adaptation** that couples outer harness evolution with inner adaptation.

---

## Introduction and Theoretical Foundation

Modern LLM-based agents operate through a **harness**: the tools they can invoke, the skills or procedures they can reuse, and the other agents they coordinate with. This harness shapes both what an agent observes and what it can do. Critically, the harness is **non-stationary** in deployed systems—it continually expands as new capabilities are introduced, updated, or retired (e.g., Salesforce's Agentforce ecosystem and OpenAI's public skills repository show steady growth over time).

The paper formalizes two distinct challenges posed by harness evolution:

1. **Retention**: As the harness expands, previously solved tasks must still be solved. The required capabilities remain present but are embedded in a larger pool, making previously successful behavior harder to recover. This leads to **harness-induced forgetting**: model parameters remain unchanged, yet competence degrades solely because the harness has expanded.

2. **Adaptation under accumulated experience**: In self-evolving systems, persistent artifacts (memories, learned skills, prompts, routing policies) are accumulated under *earlier, narrower* harnesses. As new capabilities arrive, these artifacts may become stale or misdirect execution toward behaviors that were effective previously but are no longer appropriate.

The formal framework defines a sequence of harness states:

$$H_{1:T} \triangleq (H_1, H_2, \ldots, H_T), \quad H_1 \subseteq H_2 \subseteq \cdots \subseteq H_T$$

Each task $i$ at stage $t$ has a hidden ground-truth capability set $C_i$, satisfying:
1. **Feasibility**: $C_i \subseteq H_{t_i}$ (all needed capabilities available at introduction)
2. **New-capability pressure**: $C_i \cap (H_{t_i} \setminus H_{t_i-1}) 
eq \emptyset$ (at least one newly introduced capability required)

---

## Methodology

### Construction Pipeline

The benchmark converts static agent benchmarks into evolving harness streams without creating new tasks:

**Step 1: Capability annotation.** Each task $i$ is associated with an axis-specific capability set $C_i$. For tools, annotations come from oracle tool labels; for skills and agents, they are constructed deterministically.

**Step 2: Frequency-ranked capability release.** Capabilities are ranked by frequency $f(c) = |\{i : c \in C_i\}|$ and partitioned into $T$ release buckets, so frequently used capabilities are introduced earlier (the "core"), with specialized ones released progressively toward the "long tail."

**Step 3: Cumulative harness construction.**

$$H_t := \{c \in U : r(c) \leq t\}$$

where $r(c)$ is the release stage of capability $c$.

**Step 4: Task assignment.** Each task is assigned to the earliest stage where all required capabilities are available:

$$t_i := \min\{t : C_i \subseteq H_t\} = \max_{c \in C_i} r(c)$$

### Axis-Specific Construction

- **Tools**: Uses oracle tool sets directly; the agent receives the full cumulative catalog (including distractors).
- **Skills**: Reference skills are mined from source system prompts via rule-based extraction; tasks are associated with skills via keyword matching against verifier-checked states.
- **Agents**: Tools are grouped by *owner entity* (e.g., database, filesystem, email) to form specialist agents; tasks spanning multiple entities require delegation.

### Evaluation Protocol

- **Deployment evaluation**: No persistent state; fresh system at each stage evaluated on cumulative task set $D_{\leq t, eval}$.
- **Self-evolving adaptation evaluation**: Persistent state $z_t$ is updated using adaptation split $D_{\leq t, adapt}$; evaluation on held-out split.
- Metrics: **Pass** (verifier success rate), **Score** (average verifier score), **Forward Transfer (FWT)**, and **Backward Transfer (BWT)**.

---

## Empirical Validation / Results

### Evolving Tools (Table 2)

| Setting | EOG Pass (%) | ALE Pass (%) | Overall Pass Rate |
|---|---|---|---|
| Task-specific (reference) | 26.0 | 11.6 | 24.2 |
| Deployment (cumulative) | 30.2 | 12.2 | 28.1 |
| **MemToolAgent** (best) | **38.6** | 10.1 | **35.1** |
| ReasoningBank | 36.9 | 11.1 | 33.8 |
| Meta-Harness | 35.2 | 11.1 | 32.3 |

**Key observations:**
- Broader tool exposure *improves* accuracy but raises token cost substantially.
- Adaptation gains are inconsistent: strong on EOG, weak on ALE.
- Deployment BWT is negative (harness-induced forgetting of −5.3% on EOG); adaptation can improve BWT but at the cost of negative FWT (as low as −28.5% on ALE).

### Evolving Skills (Table 4)

| Setting | EOG Pass (%) | ALE Pass (%) | Overall Pass Rate |
|---|---|---|---|
| Task-specific (reference) | 18.9 | 7.4 | 15.3 |
| Deployment (cumulative) | 18.9 | 8.3 | 15.6 |
| **GEPA** (best) | **24.1** | 7.4 | **18.8** |

**Key observations:**
- Skill expansion has little aggregate deployment effect (skills must be explicitly retrieved).
- GEPA is the clearest winner; memory-based methods are less effective than for tools.
- Skill engagement is sparse: GPT-5 invokes ~0% of offered skills, while Codex (GPT-5.5) invokes 82% in task-specific settings. GEPA raises GPT-5's invocation to 31%.

### Evolving Agents (Table 6)

| Setting | EOG Pass (%) | ALE Pass (%) | Overall Pass Rate |
|---|---|---|---|
| Task-specific (reference) | 6.5 | 5.3 | 6.2 |
| Deployment (cumulative) | 8.8 | 4.2 | 7.4 |
| **Meta-Harness** (best) | **18.5** | 3.2 | **13.9** |
| GEPA | 14.9 | 6.3 | 12.3 |

**Key observations:**
- Agent-pool expansion is strongly environment-dependent (improves EOG, degrades ALE).
- Largest adaptation gains occur here (+110.2% for Meta-Harness on EOG), but severe forgetting on ALE (deployment BWT of −34.7%).
- Adaptation improves delegation **completeness** (recall: 79% → 89%) more than selection **precision** (~90% throughout).
- Harness-induced forgetting arises mainly from **delegation drift**: forgotten tasks stop invoking previously used required agents.

### Cross-Axis Summary (Table 8)

| | Tools | Skills | Agents |
|---|---|---|---|
| Deployment expansion | Accuracy ↑, cost ↑ | Little effect | Environment-dependent |
| Max harness-induced forgetting | −5.3% | −4.0% | −34.7% |
| Best adaptation gain | +27.8% (MemToolAgent) | +27.5% (GEPA) | +110.2% (Meta-Harness) |
| Best adaptation context | Task-specific | Task-specific | Final full pool from start |
| Observed bottleneck | Extract useful tool experience | Engage the right skills | Balance delegation coverage & selectivity |

---

## Theoretical and Practical Implications

1. **Harness evolution creates a distinct retention challenge**: Negative deployment BWT occurs even with fixed model parameters, showing that prior competence degrades purely from harness expansion. The mechanism differs by axis—tools enlarge the action space, skills require retrieval and engagement, agents add coordination demands.

2. **Adaptation quality depends on the harness evolution trajectory**: Tools and skills benefit most from task-relevant exposure during adaptation, while agents improve more under a stable broader pool. This shows that inner adaptation must be designed jointly with the outer harness evolution.

3. **Retention and adaptation are separable, often conflicting objectives**: FWT and BWT trade-offs reveal that a system can appear to improve overall while becoming worse at either exploiting new capabilities or preserving prior competence. These should be measured separately throughout the harness trajectory.

4. **Practical implications for deployed systems**: Real-world platforms (e.g., Agentforce, OpenAI skills) exhibit exactly this kind of non-stationary harness. The findings suggest that:
   - Broader tool access can improve accuracy but at significant cost.
   - Skill libraries are benign at deployment but complicate inner adaptation by diluting experience.
   - Agent-pool expansion requires careful monitoring of delegation stability.

---

## Conclusion

EVOHARNESSBENCH is the first benchmark to place non-stationarity in the *externally supplied harness* rather than the task stream. The central finding is that **current agents do not yet keep pace reliably with harness evolution**—harness expansion alone can induce forgetting, self-evolving adaptation gains are inconsistent, and retention and adaptation are often in tension.

The paper motivates three future directions:

1. **Bi-level adaptation**: Inner updates optimize persistent state from current experience, while an outer objective evaluates whether those updates remain effective across harness stages—balancing adaptation to new capabilities with preservation of prior competence.

2. **Selective artifact revision**: Mechanisms to detect when persistent artifacts have become stale or harmful, and to selectively revise, retain, or discard them as the harness changes.

3. **Beyond monotonic growth**: Extending to capability replacement and retirement would introduce migration and obsolescence challenges common in real deployments.

The benchmark and code are publicly available at https://mas-orchestra.salesforceresearch.ai/evoharness/.

---

_Markdown view of https://picx.dev/p/L2X26Z, served by PicX — AI-generated visual whiteboard summaries of research papers._
