# Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild

> Agent harnesses remain substantially undertested, with less than half of LLM-dependent code covered, and HarnessTester's contract-aware test generation boosts coverage and mutation scores by up to 95% while finding 88 previously-unknown bugs.

- **Source:** [arXiv](https://arxiv.org/abs/2610.04921)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/4vikwG
- **Whiteboard:** https://picx.dev/p/4vikwG/image

## Summary

## Summary (Overview)

- This paper presents the **first empirical study** of test adequacy in real-world agent harnesses, with a novel focus on **LLM-dependent harness (LDH) code** — harness code with data or control dependencies on LLM outputs.
- The study of 10 widely-used agentic systems (e.g., OpenClaw, OpenHands, Aider) reveals that **agent harnesses remain substantially undertested**: less than half of LDH lines and branches are covered, with an average mutation score of only 33.60%.
- The authors propose **HarnessTester**, the first harness-oriented test generation technique that incorporates explicit **agent-harness contract support** to construct contract-faithful test setups.
- HarnessTester outperforms state-of-the-art general-purpose test generation baselines by **75.95%/84.76% larger line/branch coverage gains** and **69.89% larger mutation-score gains**.
- HarnessTester detects **122 real-world harness bugs** in widely-used agentic systems, including 88 previously-unknown bugs, 69 of which have been confirmed by developers.

## Introduction and Theoretical Foundation

LLM-based agentic systems are emerging as a new software paradigm, composed of probabilistic backbone LLMs and a surrounding **harness** — the operational software infrastructure providing execution environment management, context/memory management, guardrails, workflow orchestration, and tool interfaces. Harnesses are large codebases (e.g., OpenClaw's harness exceeds 196,000 lines) and substantial performance gains can be achieved by optimizing harnesses alone.

Despite the growing adoption of agentic systems, **how adequately harness implementations are tested remains largely unexplored**. Existing benchmarks evaluate agents as black boxes via end-to-end task success rates, without examining how thoroughly internal harness implementations are exercised.

The paper introduces the concept of **LLM-dependent harness (LDH) code**, formally defined as:

$$
LDH(A) = \{h \in \mathcal{H}(A) | \exists s \in \mathcal{S}_{LLM}. s \rightarrow_{d} h\},
$$

where $\mathcal{H}(A) \subseteq \mathrm{Stmt}(A)$ denotes the set of harness statements, $\mathcal{S}_{LLM}$ is the set of LLM invocation statements, and $s \rightarrow_{d} h$ denotes that statement $h$ is directly or transitively dependent on statement $s$.

LDH code deserves particular testing attention because it processes nondeterministic LLM outputs spanning a vast value space, and errors here can propagate through agent workflows causing incorrect behavior.

## Methodology

### Empirical Study Setup
- **LDH extraction**: An interprocedural harness analysis tool models LLM invocations as sources and propagates data/control dependency facts throughout the program until a fixed point is reached, covering assignments, field/container accesses, parameter/return flows, and taint-preserving transformations.
- **Studied agents**: 10 representative open-source agentic systems with ≥10,000 GitHub stars, covering diverse domains (coding, research, general-purpose) and both Python and TypeScript.
- **Metrics**: Line/branch coverage (project-wide and LDH-wide) and mutation scores (using mutmut and StrykerJS).

### HarnessTester Design
HarnessTester follows a coverage-driven test generation paradigm with two key innovations:

1. **Agent-Harness Contract Guidance**: Prompt-level principles guiding the LLM to preserve key contracts when constructing tests, including:
   - Maintaining expected runtime representations of provider responses, tool actions, and observations
   - Constructing agent state or interaction history required to reach target behavior
   - Respecting cross-component invocation contracts (synchronous/asynchronous conventions, return shapes)
   - Installing contract-faithful test doubles at actual binding points

2. **Agent-Harness Contract Context Retriever**: A callable tool that retrieves contract-relevant code context on demand (constructors, fixtures, collaborators, mock targets) using AST-based matching, helping the LLM construct contract-faithful test setups.

## Empirical Validation / Results

### RQ1: Overall Test Adequacy

**Table 2. Overall test adequacy (key results)**

| Agent | LDH Line Coverage | LDH Branch Coverage | Project Line Coverage | Mutation Score |
|-------|-------------------|---------------------|----------------------|----------------|
| OpenHands | 54.91% | 53.42% | 66.65% | 41.66% |
| Aider | 53.87% | 48.93% | 58.65% | 35.10% |
| SWE-agent | 44.81% | 44.15% | 58.51% | 24.11% |
| PR-Agent | 27.09% | 25.76% | 37.88% | 16.91% |
| GPT Researcher | 10.67% | 11.49% | 25.75% | 4.47% |
| Browser Use | 28.30% | 25.48% | 37.51% | 25.23% |
| RD-Agent | 5.39% | 6.73% | 20.82% | 3.58% |
| OpenClaw | 56.73% | 50.55% | 59.74% | 53.18% |
| Roo Code | 68.14% | 62.18% | 57.33% | 60.86% |
| Kimi Code | 92.83% | 83.09% | 71.98% | 70.94% |
| **Average** | **44.27%** | **41.18%** | **49.48%** | **33.60%** |

Key findings:
- LDH represents 22.35% of lines but 31.71% of branches, showing a high concentration of decision logic
- Existing tests cover less than half of LDH lines and branches on average
- Mutation scores average only 33.60%, meaning ~two-thirds of injected faults go undetected

### RQ2: Semantic Test Adequacy

Analysis of 700 sampled LDH branch predicates revealed three semantic dimensions:

**Table 3. Taxonomy of semantic patterns in LDH branch predicates**

| Component Intention | Share | Judgment Pattern | Share | LLM Value Form | Share |
|---------------------|-------|------------------|-------|----------------|-------|
| Provider I/O Normalization | 29.14% | Presence/Shape Check | 47.14% | Generated Content | 25.00% |
| Tool Action Handling | 23.57% | Value/Pattern Matching | 33.29% | Provider Response Signal | 6.57% |
| Agent Orchestration Control | 14.00% | Size/Threshold Check | 15.29% | Tool Action Payload | 21.71% |
| Context Memory Management | 13.57% | Policy Guard | 4.29% | Agent Context State | 21.71% |
| Observation Feedback Handling | 12.43% | | | Observation Feedback Signal | 25.00% |
| Error Recovery Handling | 7.29% | | | | |

Coverage analysis showed that only **25.57% of LDH branch predicates are fully covered**, 24.57% partially covered, and 49.86% entirely uncovered. Predicates with explicit schemas or bounded value spaces (tool actions, policy guards) are better covered, while those involving dynamic execution contexts or open-ended value spaces (orchestration, observation feedback) show larger gaps.

### RQ3: Test Adequacy Improvement

HarnessTester was compared against four state-of-the-art LLM-based test generation baselines (CoverUp, TestTailor for Python; Qodo Cover, TestPilot 2 for TypeScript) with a 120-minute budget per project.

Key results:
- HarnessTester achieves **25.84 and 25.41 percentage-point gains** in LDH line/branch coverage (macro-averaged)
- Compared to the strongest baseline: **82.27%/94.36% larger gains** (scoped setting), **95.29%/109.22% larger** (general setting)
- The contract-agnostic variant shows **96.51% lower line-coverage gains** and **96.24% lower branch-coverage gains**, confirming the importance of contract support
- HarnessTester achieves the highest mutation-score gain on every studied agent

### RQ4: Real-World Harness Bug Detection

**RQ4.a (Historical bugs)**: On H-Bench (250 historical harness bugs), HarnessTester detected **43 bugs (17.20%)** versus only **2 (0.80%)** for the best baseline, at a cost of $1.80 per revealed bug versus $10.84 for the baseline.

**RQ4.b (Previously-unknown bugs)**:

**Table 5. Detected real-world bugs**

| Project | New | Known | Confirmed | Fix | Cost$ |
|---------|-----|-------|-----------|-----|-------|
| OpenHands Canvas | 5 | 0 | 3 | 3 | 7.38 |
| OpenHands SDK | 10 | 3 | 10 | 9 | 7.68 |
| Aider | 9 | 4 | 9 | 9 | 6.69 |
| SWE-agent | 4 | 3 | 4 | 4 | 6.71 |
| PR-Agent | 5 | 2 | 5 | 5 | 9.16 |
| GPT Researcher | 6 | 1 | 6 | 6 | 5.80 |
| Browser Use | 15 | 8 | 13 | 12 | 7.85 |
| RD-Agent | 11 | 5 | 9 | 8 | 8.23 |
| OpenClaw | 9 | 6 | 9 | 8 | 7.15 |
| Roo Code | 11 | 0 | - | - | 7.26 |
| Kimi Code | 3 | 2 | 1 | 0 | 7.73 |
| **Total** | **88** | **34** | **69** | **64** | **81.64** |

A notable example is a **sensitive-data leakage bug in Browser Use** where a nested list within a typed action parameter bypassed redaction, leaving synthetic credentials unredacted in saved history — a bug that required schema-valid inputs embedded in valid action/history object chains to reveal.

## Theoretical and Practical Implications

- **For testing practice**: The study reveals that current agent testing practices cannot fully mitigate reliability risks from LLM–harness interactions. High-adequacy LDH testing requires contract-faithful setups satisfying expected data structures, runtime states, and interaction protocols — a challenge not addressed by general-purpose test generation.

- **For tooling**: HarnessTester demonstrates that explicitly encoding agent-harness contracts (via prompt guidance and context retrieval) substantially improves test adequacy and bug detection, suggesting a new direction for agent-specific testing tools.

- **For practitioners**: The findings suggest that agent developers should invest in reusable test utilities that encode harness contracts (like Kimi Code's approach), enabling construction of both contract-faithful and diverse execution contexts.

- **Cost-effectiveness**: HarnessTester reveals bugs at $1.80 per bug versus $10.84 for baselines, demonstrating practical deployability despite higher absolute costs.

## Conclusion

This work presents the first empirical study of test adequacy in real-world agent harnesses, revealing substantial gaps in exercising LLM-dependent harness code — particularly where testing requires dynamic execution contexts, implicit input spaces, and cross-component interactions. The proposed **HarnessTester** incorporates agent-harness contract support to construct contract-faithful tests, substantially outperforming general-purpose test generation in coverage and mutation-score gains. Its detection of 122 real-world harness bugs (88 previously-unknown, 69 confirmed by developers) highlights the practical value of harness-oriented test generation for improving the reliability of agentic systems.

**Future directions** include extending the approach to more agentic systems, exploring additional contract types, and integrating harness-oriented testing into continuous integration pipelines for agent development.

---

_Markdown view of https://picx.dev/p/4vikwG, served by PicX — AI-generated visual whiteboard summaries of research papers._
