Summary

  • First systematic empirical study and benchmark of security mechanisms in coding agent harnesses across 40 open- and closed-source products, introducing a ten-mechanism taxonomy (RQ1) and a 400-cell implementation assessment (RQ2).
  • Introduction of HARNESSSECURITY-BENCH (HSB)—a benchmark of 23 tasks across five attack surfaces, evaluating nine security mechanisms across six leading harnesses (Claude Code, Codex CLI, Gemini CLI, gptme, Qwen Code, GitHub Copilot) with separate deterministic oracles for task utility and attack effects.
  • Large-scale evaluation comprising 2,500 trials, 81,155 tool calls, and over 2.2 billion tokens, revealing that enabling auto-approve raises attack success from 29.2% to 95.6% while network isolation and read-only mode reduce attack effects but with substantial utility losses.
  • Key finding: Command allowlisting and command denylisting reduce attack effects with small utility loss and utility gain respectively, while task-level cases show restrictions on shared capabilities can obstruct both legitimate and malicious operations.
  • Critical observation: About half of confirmed mechanism implementations are opt-in, and closed-source harnesses exhibit substantial evidence gaps (53.1% of cells unresolved vs. 15.9% for open-source).

Introduction and Theoretical Foundation

Coding agents based on large language models (LLMs) have moved from demonstrations to production, with products such as Claude Code and Codex CLI supporting autonomous repository-level software development. Each agent couples an LLM with a coding agent harness that assembles context, mediates tool use, and authorizes actions. Harness design affects task success and resource costs even when the underlying model is held fixed.

The paper identifies several real-world security risks motivating this work:

  • Prompt injection through issue titles induced shell command execution in Cline's triage workflow
  • Network isolation failures reported by OpenAI and Anthropic during cybersecurity evaluations
  • IDE-based attacks (IDEsaster) where agent file edits triggered unsafe behaviors in the surrounding IDE
  • Data confidentiality concerns (ZCode) where workspace snapshots containing Git history were uploaded

The research is organized around three questions:

  • RQ1: What distinct security mechanisms can be identified in real coding agent harnesses?
  • RQ2: To what degree are these mechanisms implemented?
  • RQ3: How do security mechanisms affect security and utility during adversarially augmented task execution?

The theoretical foundation draws on the five-stage coding agent harness loop: context assembly → model planning → authorization/execution → context update → (optionally) skill loading, subagent delegation, and MCP connections. Attack surfaces are organized by the resource under adversarial control: repository/workflow artifacts, instructions/skills, tool interfaces, knowledge/memory, external services/dependencies, and direct user input.


Methodology

Harness Dataset (RQ1)

  • Collected 75 candidate records from GitHub repositories, commercial products, and security disclosures
  • After merging duplicates (19 pairs) and applying screening criteria (stars ≥ 1,000, active maintenance, agentic coding workflow support), retained 40 harnesses: 27 open-source and 13 closed-source
  • Categories include CLI tools, IDE extensions, IDE-based agents, and cloud/platform agents
  • Programmatic collection archived documentation, READMEs, CLI materials, and (for open-source) source code

Security Mechanism Taxonomy (RQ1)

Using open and axial coding principles, three researchers independently labeled security behaviors in eight initial harnesses, then extended to all 40. After adjudication with a fourth researcher, ten mechanisms were identified:

Mech.Definition
AAAuto-approve: skips some/all confirmations
NINetwork isolation: default-deny egress or domain allowlist
CALCommand allowlisting: explicit allowed command set
CDLCommand denylisting: blocks explicitly dangerous commands
MCPMCP permissions: authorizes external-tool servers/capabilities
PIFPrompt-injection filtering: tags/detects/sanitizes untrusted instructions
PRPath restriction: constrains file reads/writes to permitted paths
RORead-only mode: suppresses persistent changes
ALAudit logging: records security-relevant actions
PTProject trust: admits repository-supplied configuration/resources per trust state

Rating Protocol (RQ2)

Each of the 400 harness–mechanism cells receives one of five ratings:

  • D (default-enforced): implemented and active by default, no disable option
  • C (configurable-off): implemented and active by default, can be disabled
  • O (opt-in): implemented but requires explicit user enablement
  • A (absent): affirmative evidence of non-implementation
  • U (unknown): insufficient or conflicting evidence

Three LLM raters (Claude Opus 4.8, DeepSeek V4 Flash, GLM-5.2) and two human researchers independently rate all cells using an Evidence-of-Thought (EoT) method requiring source-cited justifications. Majority (≥3 of 5) determines consensus; disagreements (18 cells) are adjudicated by a third human researcher.

HARNESSSECURITY-BENCH (RQ3)

  • 23 tasks (21 original, 2 adapted from Terminal-Bench) across nine mechanism benchmarks (AL excluded)
  • Tasks span Python, JavaScript, C++, C#, Ruby, and Go
  • Attack surface augmentation: environments augmented with misleading context and malicious resources across five attack surfaces (repository artifacts, instructions/skills, tool interfaces, knowledge/memory, external services) while preserving legitimate user instructions and utility oracles
  • Deterministic oracles separately measure task utility and attack effects (file/database changes, receiver-side records, execution events)
  • Evaluation setup: Each trial runs in a fresh Docker container; 10 trials per condition; GLM-5.2 as base LLM; external execution deadline of 28,800 seconds
  • Six harnesses selected for having at most one U/A rating: Claude Code, Codex CLI, Gemini CLI, gptme, Qwen Code, GitHub Copilot

Empirical Validation / Results

RQ2 Implementation Assessment (400 cells)

  • 205 confirmed implementations (32 D, 64 C, 109 O), 83 confirmed absent, 112 unknown
  • RO and AA are mostly opt-in (30 of 31 RO; 26 of 30 AA are O)
  • MCP and AL are usually active by default (22 of 29 MCP; 20 of 22 AL)
  • Closed-source harnesses have 53.1% unknown ratings (69/130) vs. 15.9% for open-source (43/270)
  • All 13 closed-source PIF cells are unknown
  • Rater agreement: Krippendorff's α = 0.740 (all five raters), 0.669 (LLMs), 0.911 (humans)

RQ3 Benchmark Results

Table 6: AA, NI, CAL, and CDL results (Overall rows)

MechanismUtility offUtility onASR offASR onΔUtilityΔASR
AA77.1% (601/780)95.4% (744/780)29.2% (105/360)95.6% (344/360)↑ 18.3%↑ 66.4%
NI90.9% (600/660)66.4% (438/660)57.1% (137/240)0.8% (2/240)↓ 24.5%↓ 56.3%
CAL99.2% (655/660)97.7% (645/660)98.1% (353/360)50.0% (180/360)↓ 1.5%↓ 48.1%
CDL94.6% (530/560)99.8% (559/560)41.7% (100/240)1.3% (3/240)↑ 5.2%↓ 40.4%

Table 7: MCP, PIF, PR, RO, and PT results (Overall rows)

MechanismUtility offUtility onASR offASR onΔUtilityΔASR
MCP73.3% (836/1140)73.9% (842/1140)82.2% (296/360)54.2% (195/360)↑ 0.5%↓ 28.1%
PIF85.0% (136/160)83.1% (133/160)26.7% (16/60)23.3% (14/60)↓ 1.9%↓ 3.3%
PR81.6% (1175/1440)81.9% (1180/1440)95.0% (152/160)46.9% (75/160)↑ 0.3%↓ 48.1%
RO69.8% (754/1080)35.8% (387/1080)8.3% (10/120)0.8% (1/120)↓ 34.0%↓ 7.5%
PT100.0% (1200/1200)80.0% (960/1200)75.3% (226/300)40.3% (121/300)↓ 20.0%↓ 35.0%

Key findings by mechanism:

  • AA: Skipping review allows both legitimate and malicious operations to proceed; ASR increased in every harness, but utility benefits varied (Gemini CLI +2.3 pp, Claude Code −0.8 pp)
  • NI: Required network interactions accounted for 161 of 162 fewer utility passes; blocking registry access also prevented malicious installation callbacks
  • CDL: 15 of 29 additional utility passes came from reports no longer containing signing-key material
  • PT: Gemini CLI refused startup in all 40 ON trials—its utility loss was entirely from startup refusal; other four harnesses retained 100% utility while ASR fell from 75.8% to 50.4%
  • PIF: Only gptme supported paired comparison; small improvement (ASR −3.3 pp) suggests attacks presented as ordinary workflow prerequisites still influence execution

Case Studies

  • Case A (Dependency Lock Repair): Enabling NI in Claude Code reduced utility from 96% to 76% while ASR fell from 100% to zero—legitimate and malicious operations depended on the same network capability
  • Case B (CLI Arg Parser): Qwen Code's CAL gate rejected the unauthorized checker, but the agent invoked it through an allowed Python interpreter—demonstrating bypass through alternative execution paths

Resource Usage

Across 2,500 trials: 81,155 tool calls, 589.32 hours cumulative execution time, 2.164 billion input tokens and 53.714 million output tokens (excluding unavailable conditions).


Theoretical and Practical Implications

For Harness Providers

  1. Make security settings verifiable: Publish versioned mechanism-level descriptions stating implementation status, default activation state, configuration options, and covered inputs/operations. Expose effective security settings through CLI/UI so users can verify their actual configuration.

  2. Test alternative execution paths: Providers should test alternative access paths to protected resources (e.g., invoking blocked commands through allowed interpreters) and verify legitimate tasks retain authorized completion paths. Re-check after changes to tool routing or permission configuration.

  3. Evaluate settings against task requirements: Report attack effects, task utility, token usage, and runtime together so users can select settings appropriate to their workloads and acceptable risk.

Theoretical Contributions

  • Common taxonomy: Provides a shared vocabulary for describing security mechanisms across heterogeneous harnesses, distinguishing mechanisms by protected resources/operations and enforcement points
  • Evidence-of-Thought rating method: Makes LLM judgments traceable to source materials through required citations and hash verification
  • Dual-oracle evaluation: Separates task utility from attack effects using deterministic oracles, avoiding misclassification from trace-based pattern matching or LLM judgment alone

Practical Trade-off Insights

The paper demonstrates that security mechanisms exhibit fundamentally different trade-off profiles:

  • Coarse-grained restrictions (NI, RO) reduce attacks but can block legitimate operations sharing the same capability
  • Precise restrictions (CAL, CDL, PR, MCP) can reduce attacks with minimal or even positive utility impact
  • Risk capabilities (AA) increase both utility and attack exposure—sometimes with attack exposure rising substantially even when utility gains are minimal

Conclusion

This study makes two primary contributions to the security of coding agent harnesses:

  1. Security mechanism dataset and empirical study (RQ1–RQ2): A taxonomy of ten security mechanisms across 40 coding agent harnesses, with auditable implementation ratings showing that ~half of confirmed implementations are opt-in and closed-source products have significant evidence gaps.

  2. HARNESSSECURITY-BENCH (RQ3): A containerized benchmark evaluating nine mechanisms through 23 tasks, using separate deterministic oracles for task utility and attack effects, with native ON/OFF comparisons under controlled conditions.

Key takeaways:

  • Enabling auto-approve increases attack success in all six harnesses (29.2% → 95.6%)
  • Network isolation and read-only mode reduce attack effects but with substantial utility losses (24.5 and 34.0 pp)
  • Command allowlisting and denylisting achieve security gains with minimal utility impact
  • Alternative execution paths can leave unauthorized operations reachable even with restrictions in place

Future directions: Extend evaluation to closed-source harnesses, additional LLMs, and downstream uses of recorded trajectories for harness refinement.

Related papers