Summary (Overview)

  • This paper presents the first empirical study of test adequacy in real-world agent harnesses, with a novel focus on LLM-dependent harness (LDH) code — harness code with data or control dependencies on LLM outputs.
  • The study of 10 widely-used agentic systems (e.g., OpenClaw, OpenHands, Aider) reveals that agent harnesses remain substantially undertested: less than half of LDH lines and branches are covered, with an average mutation score of only 33.60%.
  • The authors propose HarnessTester, the first harness-oriented test generation technique that incorporates explicit agent-harness contract support to construct contract-faithful test setups.
  • HarnessTester outperforms state-of-the-art general-purpose test generation baselines by 75.95%/84.76% larger line/branch coverage gains and 69.89% larger mutation-score gains.
  • HarnessTester detects 122 real-world harness bugs in widely-used agentic systems, including 88 previously-unknown bugs, 69 of which have been confirmed by developers.

Introduction and Theoretical Foundation

LLM-based agentic systems are emerging as a new software paradigm, composed of probabilistic backbone LLMs and a surrounding harness — the operational software infrastructure providing execution environment management, context/memory management, guardrails, workflow orchestration, and tool interfaces. Harnesses are large codebases (e.g., OpenClaw's harness exceeds 196,000 lines) and substantial performance gains can be achieved by optimizing harnesses alone.

Despite the growing adoption of agentic systems, how adequately harness implementations are tested remains largely unexplored. Existing benchmarks evaluate agents as black boxes via end-to-end task success rates, without examining how thoroughly internal harness implementations are exercised.

The paper introduces the concept of LLM-dependent harness (LDH) code, formally defined as:

LDH(A)={h∈H(A)∣∃s∈SLLM.s→dh},LDH(A) = \{h \in \mathcal{H}(A) | \exists s \in \mathcal{S}_{LLM}. s \rightarrow_{d} h\},

where H(A)⊆Stmt(A)\mathcal{H}(A) \subseteq \mathrm{Stmt}(A) denotes the set of harness statements, SLLM\mathcal{S}_{LLM} is the set of LLM invocation statements, and s→dhs \rightarrow_{d} h denotes that statement hh is directly or transitively dependent on statement ss.

LDH code deserves particular testing attention because it processes nondeterministic LLM outputs spanning a vast value space, and errors here can propagate through agent workflows causing incorrect behavior.

Methodology

Empirical Study Setup

  • LDH extraction: An interprocedural harness analysis tool models LLM invocations as sources and propagates data/control dependency facts throughout the program until a fixed point is reached, covering assignments, field/container accesses, parameter/return flows, and taint-preserving transformations.
  • Studied agents: 10 representative open-source agentic systems with ≥10,000 GitHub stars, covering diverse domains (coding, research, general-purpose) and both Python and TypeScript.
  • Metrics: Line/branch coverage (project-wide and LDH-wide) and mutation scores (using mutmut and StrykerJS).

HarnessTester Design

HarnessTester follows a coverage-driven test generation paradigm with two key innovations:

  1. Agent-Harness Contract Guidance: Prompt-level principles guiding the LLM to preserve key contracts when constructing tests, including:

    • Maintaining expected runtime representations of provider responses, tool actions, and observations
    • Constructing agent state or interaction history required to reach target behavior
    • Respecting cross-component invocation contracts (synchronous/asynchronous conventions, return shapes)
    • Installing contract-faithful test doubles at actual binding points
  2. Agent-Harness Contract Context Retriever: A callable tool that retrieves contract-relevant code context on demand (constructors, fixtures, collaborators, mock targets) using AST-based matching, helping the LLM construct contract-faithful test setups.

Empirical Validation / Results

RQ1: Overall Test Adequacy

Table 2. Overall test adequacy (key results)

AgentLDH Line CoverageLDH Branch CoverageProject Line CoverageMutation Score
OpenHands54.91%53.42%66.65%41.66%
Aider53.87%48.93%58.65%35.10%
SWE-agent44.81%44.15%58.51%24.11%
PR-Agent27.09%25.76%37.88%16.91%
GPT Researcher10.67%11.49%25.75%4.47%
Browser Use28.30%25.48%37.51%25.23%
RD-Agent5.39%6.73%20.82%3.58%
OpenClaw56.73%50.55%59.74%53.18%
Roo Code68.14%62.18%57.33%60.86%
Kimi Code92.83%83.09%71.98%70.94%
Average44.27%41.18%49.48%33.60%

Key findings:

  • LDH represents 22.35% of lines but 31.71% of branches, showing a high concentration of decision logic
  • Existing tests cover less than half of LDH lines and branches on average
  • Mutation scores average only 33.60%, meaning ~two-thirds of injected faults go undetected

RQ2: Semantic Test Adequacy

Analysis of 700 sampled LDH branch predicates revealed three semantic dimensions:

Table 3. Taxonomy of semantic patterns in LDH branch predicates

Component IntentionShareJudgment PatternShareLLM Value FormShare
Provider I/O Normalization29.14%Presence/Shape Check47.14%Generated Content25.00%
Tool Action Handling23.57%Value/Pattern Matching33.29%Provider Response Signal6.57%
Agent Orchestration Control14.00%Size/Threshold Check15.29%Tool Action Payload21.71%
Context Memory Management13.57%Policy Guard4.29%Agent Context State21.71%
Observation Feedback Handling12.43%Observation Feedback Signal25.00%
Error Recovery Handling7.29%

Coverage analysis showed that only 25.57% of LDH branch predicates are fully covered, 24.57% partially covered, and 49.86% entirely uncovered. Predicates with explicit schemas or bounded value spaces (tool actions, policy guards) are better covered, while those involving dynamic execution contexts or open-ended value spaces (orchestration, observation feedback) show larger gaps.

RQ3: Test Adequacy Improvement

HarnessTester was compared against four state-of-the-art LLM-based test generation baselines (CoverUp, TestTailor for Python; Qodo Cover, TestPilot 2 for TypeScript) with a 120-minute budget per project.

Key results:

  • HarnessTester achieves 25.84 and 25.41 percentage-point gains in LDH line/branch coverage (macro-averaged)
  • Compared to the strongest baseline: 82.27%/94.36% larger gains (scoped setting), 95.29%/109.22% larger (general setting)
  • The contract-agnostic variant shows 96.51% lower line-coverage gains and 96.24% lower branch-coverage gains, confirming the importance of contract support
  • HarnessTester achieves the highest mutation-score gain on every studied agent

RQ4: Real-World Harness Bug Detection

RQ4.a (Historical bugs): On H-Bench (250 historical harness bugs), HarnessTester detected 43 bugs (17.20%) versus only 2 (0.80%) for the best baseline, at a cost of 1.80perrevealedbugversus1.80 per revealed bug versus 10.84 for the baseline.

RQ4.b (Previously-unknown bugs):

Table 5. Detected real-world bugs

ProjectNewKnownConfirmedFixCost$
OpenHands Canvas50337.38
OpenHands SDK1031097.68
Aider94996.69
SWE-agent43446.71
PR-Agent52559.16
GPT Researcher61665.80
Browser Use15813127.85
RD-Agent115988.23
OpenClaw96987.15
Roo Code110--7.26
Kimi Code32107.73
Total8834696481.64

A notable example is a sensitive-data leakage bug in Browser Use where a nested list within a typed action parameter bypassed redaction, leaving synthetic credentials unredacted in saved history — a bug that required schema-valid inputs embedded in valid action/history object chains to reveal.

Theoretical and Practical Implications

  • For testing practice: The study reveals that current agent testing practices cannot fully mitigate reliability risks from LLM–harness interactions. High-adequacy LDH testing requires contract-faithful setups satisfying expected data structures, runtime states, and interaction protocols — a challenge not addressed by general-purpose test generation.

  • For tooling: HarnessTester demonstrates that explicitly encoding agent-harness contracts (via prompt guidance and context retrieval) substantially improves test adequacy and bug detection, suggesting a new direction for agent-specific testing tools.

  • For practitioners: The findings suggest that agent developers should invest in reusable test utilities that encode harness contracts (like Kimi Code's approach), enabling construction of both contract-faithful and diverse execution contexts.

  • Cost-effectiveness: HarnessTester reveals bugs at 1.80perbugversus1.80 per bug versus 10.84 for baselines, demonstrating practical deployability despite higher absolute costs.

Conclusion

This work presents the first empirical study of test adequacy in real-world agent harnesses, revealing substantial gaps in exercising LLM-dependent harness code — particularly where testing requires dynamic execution contexts, implicit input spaces, and cross-component interactions. The proposed HarnessTester incorporates agent-harness contract support to construct contract-faithful tests, substantially outperforming general-purpose test generation in coverage and mutation-score gains. Its detection of 122 real-world harness bugs (88 previously-unknown, 69 confirmed by developers) highlights the practical value of harness-oriented test generation for improving the reliability of agentic systems.

Future directions include extending the approach to more agentic systems, exploring additional contract types, and integrating harness-oriented testing into continuous integration pipelines for agent development.

Related papers