Full text not available for this paper

Summary (Overview)

  • EVOHARNESSBENCH is a new benchmark for evaluating LLM-based agents under harness evolution—the controlled, cumulative growth of externally supplied tools, skills, and specialist agents.
  • The benchmark contains 17 multi-stage harness streams (3–6 stages each) with 802 unique tasks, 520 tools, 42 latent reference skills, and 62 specialist agents, constructed deterministically from verifier-based benchmarks (EnterpriseOps-Gym and Agents' Last Exam).
  • Two evaluation settings are introduced: deployment evaluation (no persistent state; isolates harness-induced forgetting) and self-evolving adaptation evaluation (persistent artifacts carried across stages).
  • Key findings reveal: (1) harness expansion alone can degrade performance on previously solved tasks (harness-induced forgetting, up to −46.4% for agents); (2) self-evolving adaptation gains are inconsistent across stages, axes, and environments; (3) retention and adaptation are often in tension.
  • The work establishes harness evolution as a distinct challenge, motivating bi-level adaptation that couples outer harness evolution with inner adaptation.

Introduction and Theoretical Foundation

Modern LLM-based agents operate through a harness: the tools they can invoke, the skills or procedures they can reuse, and the other agents they coordinate with. This harness shapes both what an agent observes and what it can do. Critically, the harness is non-stationary in deployed systems—it continually expands as new capabilities are introduced, updated, or retired (e.g., Salesforce's Agentforce ecosystem and OpenAI's public skills repository show steady growth over time).

The paper formalizes two distinct challenges posed by harness evolution:

  1. Retention: As the harness expands, previously solved tasks must still be solved. The required capabilities remain present but are embedded in a larger pool, making previously successful behavior harder to recover. This leads to harness-induced forgetting: model parameters remain unchanged, yet competence degrades solely because the harness has expanded.

  2. Adaptation under accumulated experience: In self-evolving systems, persistent artifacts (memories, learned skills, prompts, routing policies) are accumulated under earlier, narrower harnesses. As new capabilities arrive, these artifacts may become stale or misdirect execution toward behaviors that were effective previously but are no longer appropriate.

The formal framework defines a sequence of harness states:

H1:T≜(H1,H2,…,HT),H1⊆H2⊆⋯⊆HTH_{1:T} \triangleq (H_1, H_2, \ldots, H_T), \quad H_1 \subseteq H_2 \subseteq \cdots \subseteq H_T

Each task ii at stage tt has a hidden ground-truth capability set CiC_i, satisfying:

  1. Feasibility: Ci⊆HtiC_i \subseteq H_{t_i} (all needed capabilities available at introduction)
  2. New-capability pressure: Ci∩(Hti∖Hti−1)eq∅C_i \cap (H_{t_i} \setminus H_{t_i-1}) eq \emptyset (at least one newly introduced capability required)

Methodology

Construction Pipeline

The benchmark converts static agent benchmarks into evolving harness streams without creating new tasks:

Step 1: Capability annotation. Each task ii is associated with an axis-specific capability set CiC_i. For tools, annotations come from oracle tool labels; for skills and agents, they are constructed deterministically.

Step 2: Frequency-ranked capability release. Capabilities are ranked by frequency f(c)=∣{i:c∈Ci}∣f(c) = |\{i : c \in C_i\}| and partitioned into TT release buckets, so frequently used capabilities are introduced earlier (the "core"), with specialized ones released progressively toward the "long tail."

Step 3: Cumulative harness construction.

Ht:={c∈U:r(c)≤t}H_t := \{c \in U : r(c) \leq t\}

where r(c)r(c) is the release stage of capability cc.

Step 4: Task assignment. Each task is assigned to the earliest stage where all required capabilities are available:

ti:=min⁡{t:Ci⊆Ht}=max⁡c∈Cir(c)t_i := \min\{t : C_i \subseteq H_t\} = \max_{c \in C_i} r(c)

Axis-Specific Construction

  • Tools: Uses oracle tool sets directly; the agent receives the full cumulative catalog (including distractors).
  • Skills: Reference skills are mined from source system prompts via rule-based extraction; tasks are associated with skills via keyword matching against verifier-checked states.
  • Agents: Tools are grouped by owner entity (e.g., database, filesystem, email) to form specialist agents; tasks spanning multiple entities require delegation.

Evaluation Protocol

  • Deployment evaluation: No persistent state; fresh system at each stage evaluated on cumulative task set D≤t,evalD_{\leq t, eval}.
  • Self-evolving adaptation evaluation: Persistent state ztz_t is updated using adaptation split D≤t,adaptD_{\leq t, adapt}; evaluation on held-out split.
  • Metrics: Pass (verifier success rate), Score (average verifier score), Forward Transfer (FWT), and Backward Transfer (BWT).

Empirical Validation / Results

Evolving Tools (Table 2)

SettingEOG Pass (%)ALE Pass (%)Overall Pass Rate
Task-specific (reference)26.011.624.2
Deployment (cumulative)30.212.228.1
MemToolAgent (best)38.610.135.1
ReasoningBank36.911.133.8
Meta-Harness35.211.132.3

Key observations:

  • Broader tool exposure improves accuracy but raises token cost substantially.
  • Adaptation gains are inconsistent: strong on EOG, weak on ALE.
  • Deployment BWT is negative (harness-induced forgetting of −5.3% on EOG); adaptation can improve BWT but at the cost of negative FWT (as low as −28.5% on ALE).

Evolving Skills (Table 4)

SettingEOG Pass (%)ALE Pass (%)Overall Pass Rate
Task-specific (reference)18.97.415.3
Deployment (cumulative)18.98.315.6
GEPA (best)24.17.418.8

Key observations:

  • Skill expansion has little aggregate deployment effect (skills must be explicitly retrieved).
  • GEPA is the clearest winner; memory-based methods are less effective than for tools.
  • Skill engagement is sparse: GPT-5 invokes ~0% of offered skills, while Codex (GPT-5.5) invokes 82% in task-specific settings. GEPA raises GPT-5's invocation to 31%.

Evolving Agents (Table 6)

SettingEOG Pass (%)ALE Pass (%)Overall Pass Rate
Task-specific (reference)6.55.36.2
Deployment (cumulative)8.84.27.4
Meta-Harness (best)18.53.213.9
GEPA14.96.312.3

Key observations:

  • Agent-pool expansion is strongly environment-dependent (improves EOG, degrades ALE).
  • Largest adaptation gains occur here (+110.2% for Meta-Harness on EOG), but severe forgetting on ALE (deployment BWT of −34.7%).
  • Adaptation improves delegation completeness (recall: 79% → 89%) more than selection precision (~90% throughout).
  • Harness-induced forgetting arises mainly from delegation drift: forgotten tasks stop invoking previously used required agents.

Cross-Axis Summary (Table 8)

ToolsSkillsAgents
Deployment expansionAccuracy ↑, cost ↑Little effectEnvironment-dependent
Max harness-induced forgetting−5.3%−4.0%−34.7%
Best adaptation gain+27.8% (MemToolAgent)+27.5% (GEPA)+110.2% (Meta-Harness)
Best adaptation contextTask-specificTask-specificFinal full pool from start
Observed bottleneckExtract useful tool experienceEngage the right skillsBalance delegation coverage & selectivity

Theoretical and Practical Implications

  1. Harness evolution creates a distinct retention challenge: Negative deployment BWT occurs even with fixed model parameters, showing that prior competence degrades purely from harness expansion. The mechanism differs by axis—tools enlarge the action space, skills require retrieval and engagement, agents add coordination demands.

  2. Adaptation quality depends on the harness evolution trajectory: Tools and skills benefit most from task-relevant exposure during adaptation, while agents improve more under a stable broader pool. This shows that inner adaptation must be designed jointly with the outer harness evolution.

  3. Retention and adaptation are separable, often conflicting objectives: FWT and BWT trade-offs reveal that a system can appear to improve overall while becoming worse at either exploiting new capabilities or preserving prior competence. These should be measured separately throughout the harness trajectory.

  4. Practical implications for deployed systems: Real-world platforms (e.g., Agentforce, OpenAI skills) exhibit exactly this kind of non-stationary harness. The findings suggest that:

    • Broader tool access can improve accuracy but at significant cost.
    • Skill libraries are benign at deployment but complicate inner adaptation by diluting experience.
    • Agent-pool expansion requires careful monitoring of delegation stability.

Conclusion

EVOHARNESSBENCH is the first benchmark to place non-stationarity in the externally supplied harness rather than the task stream. The central finding is that current agents do not yet keep pace reliably with harness evolution—harness expansion alone can induce forgetting, self-evolving adaptation gains are inconsistent, and retention and adaptation are often in tension.

The paper motivates three future directions:

  1. Bi-level adaptation: Inner updates optimize persistent state from current experience, while an outer objective evaluates whether those updates remain effective across harness stages—balancing adaptation to new capabilities with preservation of prior competence.

  2. Selective artifact revision: Mechanisms to detect when persistent artifacts have become stale or harmful, and to selectively revise, retain, or discard them as the harness changes.

  3. Beyond monotonic growth: Extending to capability replacement and retirement would introduce migration and obsolescence challenges common in real deployments.

The benchmark and code are publicly available at https://mas-orchestra.salesforceresearch.ai/evoharness/.

Related papers