Summary (Overview)
- TESTJACK is a novel automated framework that audits coding benchmark results by evolving the evaluator rather than trusting fixed test suites, addressing a fundamental limitation of static test-based evaluation for LLM coding agents.
- The framework uses prompt-grounded differential testing: an LLM agent generates tests targeting prompt requirements a patch may violate, validates them against the ground-truth solution, and re-judges failures — ensuring every refutation is backed by a replayable witness test.
- Across 4,487 benchmark-passing trials from frontier models (Claude Opus, GPT-5.5/5.6, Gemini 3.1 Pro) on 5 benchmarks (DeepSWE, SkillsBench, SWE-Marathon, SWE-bench Verified, SWE-bench Pro), TESTJACK refutes 34.4% of reported successes, lowering resolution rates from 50.6% to 33.2%.
- A lightweight variant combining task-level analysis with randomized trial auditing and cross-trial witness sharing recovers 72.4% of refutations at 15.4% of the cost, achieving 7.3× the recall of the state-of-the-art static augmentation method SWE-ABS.
- The results reveal that even benchmarks designed specifically to avoid weak tests (DeepSWE: 43.5% refuted, SWE-Marathon: 53.6% refuted) substantially overstate agent capabilities, demonstrating that evaluators must become adaptive as agents optimize against fixed evaluators.
Introduction and Theoretical Foundation
Background and Motivation
LLM agents are rapidly transforming software engineering, with systems like Claude Code, Codex, and Copilot becoming integral to research and production. Code benchmarks have become the standard yardstick for comparing these agents, with 210 benchmarks introduced in 2024 and 316 in 2025. However, nearly all share a decades-old evaluation schema: a task counts as solved once the submission passes a test suite fixed when the benchmark was published.
The Core Problem
The authors identify a fundamental mismatch between static tests and dynamic agents:
"A fixed test suite checks only the behaviors its authors anticipated, whereas an agent observes execution feedback, adapts to it, and can pass every test without fully satisfying the task."
Two kinds of benchmark defects are identified:
- Insufficient tests: tests check less than the prompt requires, so trials that overlook required cases or game the tests still pass.
- Underspecified prompts: a behavior is only implied by the prompt, making it ambiguous whether it must be implemented; the ground truth implements it but no test checks it.
Theoretical Formalization
The paper formalizes the problem using behavior spaces. A benchmark instance consists of:
- Task specification (natural language, e.g., GitHub issue)
- Repository snapshot
- Ground-truth implementation (developer-written reference patch)
- Static test suite
Requirements and behavior spaces: Let denote task requirements, where contains explicit requirements and those that follow implicitly. For an artifact , denotes the behaviors it pins down. The key nesting assumption for well-formed tasks is:
The first inclusion reflects that tests check only part of what the prompt requires; the second reflects that the ground truth implements every required behavior plus incidental choices. Passing trials are guaranteed to agree with the task only on , so defects lie in the gap .
Witness tests: Auditing a trial amounts to searching this gap for a witness test satisfying:
Limitations of Prior Work
| Approach | Limitation |
|---|---|
| Manual audits (SWE-Bench+) | Do not scale; must be repeated for every agent, model, and run |
| Automated auditors (ABA, BenchGuard) | Audit the benchmark, not individual trials; verdicts rest on LLM judgment |
| Static test augmentation (UTBoost, SWE-ABS) | Strengthen tests before any submission is observed; catch only anticipated failures |
| Dynamic differential testing (PatchDiff) | Only 28.6% of divergent patches confirmable as incorrect; requires manual inspection |
TESTJACK occupies the automated, dynamic quadrant: it audits individual trials, derives tests from the task prompt, uses ground truth only to validate expectations, and backs every verdict with an executable witness test.
Methodology
Per-Trial Auditing via Prompt-Grounded Differential Testing
TESTJACK implements auditing with two LLM agents (audit agent and reviewer agent) operating inside the benchmark's execution environment. The audit agent proceeds in five steps:
-
Execution worlds: Build three repository copies — model world , reference world , and null world (without any patch, revealing which tests pass without a fix).
-
Requirements checklist: Decompose into atomic, testable clauses instantiating , recording whether covers each clause and whether the trial implements it.
-
Evidence collection: Examine the trial's diff along three axes:
- Ghost branches: code paths never executed by , identified via branch coverage
- Reward hacking: detectors for hard-coded values, input special-casing, test-environment sniffing, harness tampering, stubs
- General defects: specification violations, robustness/security issues, prioritized via mutation testing
-
Issue-demonstrating tests: Write tests asserting the behavior each violated clause requires; tests must fail in the model world and pass in the reference world (oracle).
-
Audit report: Assign verdicts (sound, incorrect despite passing, reward hack) with explanations of each defect, why it occurs, which clause requires correct behavior, and why fails to catch it.
The reviewer agent then re-judges these tests using as the sole authority (deliberately not using to infer requirements), labeling each as valid, questionable, or invalid. The trial is refuted if at least one valid test fails on it.
Lightweight Approach: Randomized Auditing with Witness Sharing
To reduce cost, TESTJACK offers a two-stage lightweight variant:
Stage 1 — Agentic task analysis: An agent reads , , and , identifies requirements of that does not adequately cover, and writes new tests targeting them. Tests failing on are refined using execution feedback or discarded after a refinement budget. This produces an augmented suite , built once per task.
Stage 2 — Randomized audit with cross-trial witness sharing:
- Let denote the current test suite (initialized to ); trials failing any test in are refuted; remaining trials form the active set .
- Loop: (i) sample trial uniformly from ; (ii) run per-trial audit on ; (iii) if witnesses are confirmed, add to , execute on every trial in , refuting all that fail; (iv) otherwise mark as clean and remove from .
- Terminates when is empty or budget is exhausted.
Soundness of sharing: A witness belongs to the task rather than the trial on which it was found — the first two conditions of Eq.(2) do not involve at all, and the reviewer judges them against rather than the trial. Thus, deciding whether a witness refutes another trial reduces to the third condition.
The scheme is economical because generating/validating tests is expensive while replaying is cheap — the number of expensive audits approaches the number of distinct defects rather than defective trials.
Empirical Validation / Results
Experimental Setup
Benchmarks audited (Table 1):
| Benchmark | #Tasks | Languages | Task origin | #Passing trials |
|---|---|---|---|---|
| DeepSWE | 113 | TS, Go, Python, JS, Rust | Authored from scratch | 2,848 |
| SkillsBench* | 13 | Various | Community-authored | 55 |
| SWE-Marathon‡ | 2 | Various | Expert-authored | 28 |
| SWE-bench Verified | 500 | Python | Mined PRs | 805 |
| SWE-bench Pro† | 731 | Python, Go, JS, TS | Mined commits | 751 |
| Total | 1,359 | 4,487 |
Agents: Claude Opus 4.6/4.7/4.8, GPT 5.5/5.6, Gemini 3.1 Pro. TESTJACK itself uses GPT-5.5 with high reasoning effort, capped at generate-and-validate iterations per trial.
RQ1: Per-Trial Auditing Results
Table 2: Per-trial auditing results
| Benchmark | #All | #Passing | #Refuted | Refuted (%) | Corrected rate |
|---|---|---|---|---|---|
| DeepSWE | 6,236 | 2,848 | 1,240 | 43.5 | 45.7% → 25.8% |
| SkillsBench | 85 | 55 | 10 | 18.2 | 64.7% → 52.9% |
| SWE-Marathon | 79 | 28 | 15 | 53.6 | 35.4% → 16.5% |
| SWE-bench Verified | 1,000 | 805 | 61 | 7.6 | 80.5% → 74.4% |
| SWE-bench Pro | 1,462 | 751 | 219 | 29.2 | 51.4% → 36.4% |
| Total | 8,862 | 4,487 | 1,545 | 34.4 | 50.6% → 33.2% |
Key finding: Careful verifier design does not remove the problem — the from-scratch verifiers of DeepSWE and SWE-Marathon yield the two highest refutation rates (43.5% and 53.6%).
RQ2: Lightweight Approach Results
Table 3: Lightweight approach at 3 audit rounds per task
| Benchmark | #Per-trial | #Lightweight | Recall (%) | Prec. (%) | Cost ($)‡ | Red. cost (%) |
|---|---|---|---|---|---|---|
| DeepSWE | 1,240 | 1,822 | 72.7 | 49.5 | 3,904 | 84.9% |
| SkillsBench | 10 | 31 | 80.0 | 25.8 | 85 | 78.6% |
| SWE-Marathon | 15 | 10 | 40.0 | 60.0 | 48 | 53.7% |
| Total | 1,265 | 1,863 | 72.4 | 49.2† | 4,037 | 84.6% |
The lightweight approach recovers 72.4% of per-trial refutations at 15.4% of the cost. Limitations include: shared witnesses transfer well within a failure mode but rarely across modes; uniform sampling spends rounds on clean trials (hurting SWE-Marathon with only 28 passing trials).
RQ3: Comparison with SWE-ABS
Table 4: Comparison with SWE-ABS and ablation
| Benchmark | Method | #Refuted | Refuted (%) | Recall (%) | Precision (%) |
|---|---|---|---|---|---|
| DeepSWE | SWE-ABS | 246 | 8.6 | 10.0 | 50.4 |
| TESTJACK-Task Analysis | 344 | 12.1 | 13.6 | 49.1 | |
| TESTJACK-Lightweight | 1,822 | 64.0 | 72.7 | 49.5 | |
| SkillsBench | SWE-ABS | 2 | 3.6 | 10.0 | 50.0 |
| TESTJACK-Task Analysis | 0 | 0.0 | 0.0 | - | |
| TESTJACK-Lightweight | 31 | 56.4 | 80.0 | 25.8 | |
| SWE-Marathon | SWE-ABS | 0 | 0.0 | 0.0 | - |
| TESTJACK-Task Analysis | 1 | 3.6 | 6.7 | 100.0 | |
| TESTJACK-Lightweight | 10 | 35.7 | 40.0 | 60.0 | |
| Total | SWE-ABS | 248 | 8.5 | 9.9 | 50.4 |
| TESTJACK-Task Analysis | 345 | 11.8 | 13.4 | 49.3 | |
| TESTJACK-Lightweight | 1,863 | 63.6 | 72.4 | 49.2 |
TESTJACK-Lightweight achieves 7.3× the recall of SWE-ABS at comparable precision. On SWE-Marathon, SWE-ABS refutes no trials at all while TESTJACK refutes 10 of 28. Two design differences explain the gap: (1) SWE-ABS never observes submissions so must anticipate failures, while TESTJACK targets implementation-dependent defects; (2) SWE-ABS derives expectations from the gold patch rather than the prompt, penalizing implementation choices the prompt leaves open.
RQ4: Ablation Study
Task analysis alone refutes 345 trials (13.4% recall). Adding randomized auditing with witness sharing raises recall to 72.4% — a 5.4× increase — at essentially unchanged precision (49.3% vs. 49.2%). The stages are complementary: most refutations require evidence from a concrete submission.
Theoretical and Practical Implications
-
Fundamental limitation of static evaluation: The results demonstrate that fixed test suites substantially overstate the capability of frontier coding agents. As LLMs become better at optimizing against fixed evaluators (reward hacking), those evaluators must themselves become more adaptive.
-
Prompt-grounded differential testing as a paradigm: Separating what a test should check (from the prompt) from whether its expectations are correct (certified by the ground truth) provides a principled approach to dynamic evaluation that avoids both the test-oracle problem and overfitting to ground-truth implementation details.
-
Cost-effective auditing at scale: The witness-sharing mechanism exploits the observation that trials for the same task often share failure modes, making expensive audits amortizable across many trials. This makes dynamic evaluation practical for large-scale benchmarks.
-
Benchmark design implications: Even benchmarks built specifically to avoid weak tests (DeepSWE, SWE-Marathon) show the highest refutation rates, suggesting that the problem is not merely test quality but the inherent limitation of any static verifier against adaptive agents.
-
Reference results for future work: The audit results can serve as references for future benchmark design and for recalibrating reported scores of existing benchmarks.
Conclusion
TESTJACK demonstrates that 34.4% of benchmark-passing trials from frontier coding agents violate task requirements, lowering overall resolution rates from 50.6% to 33.2% across five benchmarks. Even benchmarks designed specifically to avoid weak tests show the highest violation rates, revealing a fundamental limitation of current coding-agent evaluation.
The lightweight variant recovers 72.4% of refutations at 15.4% of the cost with 7.3× the recall of static augmentation methods, making dynamic evaluation practical at scale. Every refutation is backed by a replayable witness test that checks a prompt requirement, passes on the ground truth, and fails on the trial — providing executable evidence rather than LLM judgment alone.
Future directions implied by this work include: developing evaluators that evolve adaptively with the submissions they judge, improving the precision of lightweight auditing (particularly around underspecified prompts), and extending the framework to handle cross-failure-mode witness transfer and more efficient sampling strategies.
Related papers
- LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
LOLBENCH shows top coding agents resolve only 14% of long-horizon modular development tasks, with missing cross-module context as the dominant failure bottleneck.
- What Does a Harness Buy? Tokens, Mostly
The harness barely moves pass rate on SWE-bench Verified, matching rerun noise, but decisively sets cost up to 3x via fixed preamble token sizes.
- Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs
Hybrid language models over-rely on attention and underuse recurrent memory, but an auxiliary loss forcing recurrent routing improves QA accuracy by up to 5.2% and agentic success rates.