Summary (Overview)
- This paper presents the first empirical study of test adequacy in real-world agent harnesses, with a novel focus on LLM-dependent harness (LDH) code — harness code with data or control dependencies on LLM outputs.
- The study of 10 widely-used agentic systems (e.g., OpenClaw, OpenHands, Aider) reveals that agent harnesses remain substantially undertested: less than half of LDH lines and branches are covered, with an average mutation score of only 33.60%.
- The authors propose HarnessTester, the first harness-oriented test generation technique that incorporates explicit agent-harness contract support to construct contract-faithful test setups.
- HarnessTester outperforms state-of-the-art general-purpose test generation baselines by 75.95%/84.76% larger line/branch coverage gains and 69.89% larger mutation-score gains.
- HarnessTester detects 122 real-world harness bugs in widely-used agentic systems, including 88 previously-unknown bugs, 69 of which have been confirmed by developers.
Introduction and Theoretical Foundation
LLM-based agentic systems are emerging as a new software paradigm, composed of probabilistic backbone LLMs and a surrounding harness — the operational software infrastructure providing execution environment management, context/memory management, guardrails, workflow orchestration, and tool interfaces. Harnesses are large codebases (e.g., OpenClaw's harness exceeds 196,000 lines) and substantial performance gains can be achieved by optimizing harnesses alone.
Despite the growing adoption of agentic systems, how adequately harness implementations are tested remains largely unexplored. Existing benchmarks evaluate agents as black boxes via end-to-end task success rates, without examining how thoroughly internal harness implementations are exercised.
The paper introduces the concept of LLM-dependent harness (LDH) code, formally defined as:
where denotes the set of harness statements, is the set of LLM invocation statements, and denotes that statement is directly or transitively dependent on statement .
LDH code deserves particular testing attention because it processes nondeterministic LLM outputs spanning a vast value space, and errors here can propagate through agent workflows causing incorrect behavior.
Methodology
Empirical Study Setup
- LDH extraction: An interprocedural harness analysis tool models LLM invocations as sources and propagates data/control dependency facts throughout the program until a fixed point is reached, covering assignments, field/container accesses, parameter/return flows, and taint-preserving transformations.
- Studied agents: 10 representative open-source agentic systems with ≥10,000 GitHub stars, covering diverse domains (coding, research, general-purpose) and both Python and TypeScript.
- Metrics: Line/branch coverage (project-wide and LDH-wide) and mutation scores (using mutmut and StrykerJS).
HarnessTester Design
HarnessTester follows a coverage-driven test generation paradigm with two key innovations:
-
Agent-Harness Contract Guidance: Prompt-level principles guiding the LLM to preserve key contracts when constructing tests, including:
- Maintaining expected runtime representations of provider responses, tool actions, and observations
- Constructing agent state or interaction history required to reach target behavior
- Respecting cross-component invocation contracts (synchronous/asynchronous conventions, return shapes)
- Installing contract-faithful test doubles at actual binding points
-
Agent-Harness Contract Context Retriever: A callable tool that retrieves contract-relevant code context on demand (constructors, fixtures, collaborators, mock targets) using AST-based matching, helping the LLM construct contract-faithful test setups.
Empirical Validation / Results
RQ1: Overall Test Adequacy
Table 2. Overall test adequacy (key results)
| Agent | LDH Line Coverage | LDH Branch Coverage | Project Line Coverage | Mutation Score |
|---|---|---|---|---|
| OpenHands | 54.91% | 53.42% | 66.65% | 41.66% |
| Aider | 53.87% | 48.93% | 58.65% | 35.10% |
| SWE-agent | 44.81% | 44.15% | 58.51% | 24.11% |
| PR-Agent | 27.09% | 25.76% | 37.88% | 16.91% |
| GPT Researcher | 10.67% | 11.49% | 25.75% | 4.47% |
| Browser Use | 28.30% | 25.48% | 37.51% | 25.23% |
| RD-Agent | 5.39% | 6.73% | 20.82% | 3.58% |
| OpenClaw | 56.73% | 50.55% | 59.74% | 53.18% |
| Roo Code | 68.14% | 62.18% | 57.33% | 60.86% |
| Kimi Code | 92.83% | 83.09% | 71.98% | 70.94% |
| Average | 44.27% | 41.18% | 49.48% | 33.60% |
Key findings:
- LDH represents 22.35% of lines but 31.71% of branches, showing a high concentration of decision logic
- Existing tests cover less than half of LDH lines and branches on average
- Mutation scores average only 33.60%, meaning ~two-thirds of injected faults go undetected
RQ2: Semantic Test Adequacy
Analysis of 700 sampled LDH branch predicates revealed three semantic dimensions:
Table 3. Taxonomy of semantic patterns in LDH branch predicates
| Component Intention | Share | Judgment Pattern | Share | LLM Value Form | Share |
|---|---|---|---|---|---|
| Provider I/O Normalization | 29.14% | Presence/Shape Check | 47.14% | Generated Content | 25.00% |
| Tool Action Handling | 23.57% | Value/Pattern Matching | 33.29% | Provider Response Signal | 6.57% |
| Agent Orchestration Control | 14.00% | Size/Threshold Check | 15.29% | Tool Action Payload | 21.71% |
| Context Memory Management | 13.57% | Policy Guard | 4.29% | Agent Context State | 21.71% |
| Observation Feedback Handling | 12.43% | Observation Feedback Signal | 25.00% | ||
| Error Recovery Handling | 7.29% |
Coverage analysis showed that only 25.57% of LDH branch predicates are fully covered, 24.57% partially covered, and 49.86% entirely uncovered. Predicates with explicit schemas or bounded value spaces (tool actions, policy guards) are better covered, while those involving dynamic execution contexts or open-ended value spaces (orchestration, observation feedback) show larger gaps.
RQ3: Test Adequacy Improvement
HarnessTester was compared against four state-of-the-art LLM-based test generation baselines (CoverUp, TestTailor for Python; Qodo Cover, TestPilot 2 for TypeScript) with a 120-minute budget per project.
Key results:
- HarnessTester achieves 25.84 and 25.41 percentage-point gains in LDH line/branch coverage (macro-averaged)
- Compared to the strongest baseline: 82.27%/94.36% larger gains (scoped setting), 95.29%/109.22% larger (general setting)
- The contract-agnostic variant shows 96.51% lower line-coverage gains and 96.24% lower branch-coverage gains, confirming the importance of contract support
- HarnessTester achieves the highest mutation-score gain on every studied agent
RQ4: Real-World Harness Bug Detection
RQ4.a (Historical bugs): On H-Bench (250 historical harness bugs), HarnessTester detected 43 bugs (17.20%) versus only 2 (0.80%) for the best baseline, at a cost of 10.84 for the baseline.
RQ4.b (Previously-unknown bugs):
Table 5. Detected real-world bugs
| Project | New | Known | Confirmed | Fix | Cost$ |
|---|---|---|---|---|---|
| OpenHands Canvas | 5 | 0 | 3 | 3 | 7.38 |
| OpenHands SDK | 10 | 3 | 10 | 9 | 7.68 |
| Aider | 9 | 4 | 9 | 9 | 6.69 |
| SWE-agent | 4 | 3 | 4 | 4 | 6.71 |
| PR-Agent | 5 | 2 | 5 | 5 | 9.16 |
| GPT Researcher | 6 | 1 | 6 | 6 | 5.80 |
| Browser Use | 15 | 8 | 13 | 12 | 7.85 |
| RD-Agent | 11 | 5 | 9 | 8 | 8.23 |
| OpenClaw | 9 | 6 | 9 | 8 | 7.15 |
| Roo Code | 11 | 0 | - | - | 7.26 |
| Kimi Code | 3 | 2 | 1 | 0 | 7.73 |
| Total | 88 | 34 | 69 | 64 | 81.64 |
A notable example is a sensitive-data leakage bug in Browser Use where a nested list within a typed action parameter bypassed redaction, leaving synthetic credentials unredacted in saved history — a bug that required schema-valid inputs embedded in valid action/history object chains to reveal.
Theoretical and Practical Implications
-
For testing practice: The study reveals that current agent testing practices cannot fully mitigate reliability risks from LLM–harness interactions. High-adequacy LDH testing requires contract-faithful setups satisfying expected data structures, runtime states, and interaction protocols — a challenge not addressed by general-purpose test generation.
-
For tooling: HarnessTester demonstrates that explicitly encoding agent-harness contracts (via prompt guidance and context retrieval) substantially improves test adequacy and bug detection, suggesting a new direction for agent-specific testing tools.
-
For practitioners: The findings suggest that agent developers should invest in reusable test utilities that encode harness contracts (like Kimi Code's approach), enabling construction of both contract-faithful and diverse execution contexts.
-
Cost-effectiveness: HarnessTester reveals bugs at 10.84 for baselines, demonstrating practical deployability despite higher absolute costs.
Conclusion
This work presents the first empirical study of test adequacy in real-world agent harnesses, revealing substantial gaps in exercising LLM-dependent harness code — particularly where testing requires dynamic execution contexts, implicit input spaces, and cross-component interactions. The proposed HarnessTester incorporates agent-harness contract support to construct contract-faithful tests, substantially outperforming general-purpose test generation in coverage and mutation-score gains. Its detection of 122 real-world harness bugs (88 previously-unknown, 69 confirmed by developers) highlights the practical value of harness-oriented test generation for improving the reliability of agentic systems.
Future directions include extending the approach to more agentic systems, exploring additional contract types, and integrating harness-oriented testing into continuous integration pipelines for agent development.
Related papers
- VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
VERSE shows LLM optimizers improve by evolving their own harness, but only when execution-based verification tools are provided, boosting SWE-rebench accuracy across all baselines.
- HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?
Benchmarking 40 coding agent harnesses shows auto-approve raises attack success from 29% to 96%, while command allowlisting cuts attacks with minimal utility loss.
- The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
Self-improving LLM loops on small reused evaluation sets inflate reported gains by 13–20 points, with most proposals after the first rewrite being harmful.