Summary (Overview)

  • TESTJACK is a novel automated framework that audits coding benchmark results by evolving the evaluator rather than trusting fixed test suites, addressing a fundamental limitation of static test-based evaluation for LLM coding agents.
  • The framework uses prompt-grounded differential testing: an LLM agent generates tests targeting prompt requirements a patch may violate, validates them against the ground-truth solution, and re-judges failures — ensuring every refutation is backed by a replayable witness test.
  • Across 4,487 benchmark-passing trials from frontier models (Claude Opus, GPT-5.5/5.6, Gemini 3.1 Pro) on 5 benchmarks (DeepSWE, SkillsBench, SWE-Marathon, SWE-bench Verified, SWE-bench Pro), TESTJACK refutes 34.4% of reported successes, lowering resolution rates from 50.6% to 33.2%.
  • A lightweight variant combining task-level analysis with randomized trial auditing and cross-trial witness sharing recovers 72.4% of refutations at 15.4% of the cost, achieving 7.3× the recall of the state-of-the-art static augmentation method SWE-ABS.
  • The results reveal that even benchmarks designed specifically to avoid weak tests (DeepSWE: 43.5% refuted, SWE-Marathon: 53.6% refuted) substantially overstate agent capabilities, demonstrating that evaluators must become adaptive as agents optimize against fixed evaluators.

Introduction and Theoretical Foundation

Background and Motivation

LLM agents are rapidly transforming software engineering, with systems like Claude Code, Codex, and Copilot becoming integral to research and production. Code benchmarks have become the standard yardstick for comparing these agents, with 210 benchmarks introduced in 2024 and 316 in 2025. However, nearly all share a decades-old evaluation schema: a task counts as solved once the submission passes a test suite fixed when the benchmark was published.

The Core Problem

The authors identify a fundamental mismatch between static tests and dynamic agents:

"A fixed test suite checks only the behaviors its authors anticipated, whereas an agent observes execution feedback, adapts to it, and can pass every test without fully satisfying the task."

Two kinds of benchmark defects are identified:

  1. Insufficient tests: tests check less than the prompt requires, so trials that overlook required cases or game the tests still pass.
  2. Underspecified prompts: a behavior is only implied by the prompt, making it ambiguous whether it must be implemented; the ground truth implements it but no test checks it.

Theoretical Formalization

The paper formalizes the problem using behavior spaces. A benchmark instance consists of:

  • Task specification PP (natural language, e.g., GitHub issue)
  • Repository snapshot RR
  • Ground-truth implementation GG (developer-written reference patch)
  • Static test suite T0T_0

Requirements and behavior spaces: Let Q(P)=Qe(P)∪Qi(P)\mathcal{Q}(P) = \mathcal{Q}_e(P) \cup \mathcal{Q}_i(P) denote task requirements, where Qe(P)\mathcal{Q}_e(P) contains explicit requirements and Qi(P)\mathcal{Q}_i(P) those that follow implicitly. For an artifact XX, B(X)B(X) denotes the behaviors it pins down. The key nesting assumption for well-formed tasks is:

B(T0)⊆B(P)⊆B(G).(1)\mathcal{B}(T_0) \subseteq \mathcal{B}(P) \subseteq \mathcal{B}(G).\tag{1}

The first inclusion reflects that tests check only part of what the prompt requires; the second reflects that the ground truth implements every required behavior plus incidental choices. Passing trials are guaranteed to agree with the task only on B(T0)\mathcal{B}(T_0), so defects lie in the gap B(P)∖B(T0)\mathcal{B}(P) \setminus \mathcal{B}(T_0).

Witness tests: Auditing a trial ss amounts to searching this gap for a witness test tt satisfying:

t checks some q∈Q(P)∧t(R[G])=pass∧t(R[s])=fail.(2)t \text{ checks some } q \in \mathcal{Q}(P) \quad \land \quad t(R[G]) = \text{pass} \quad \land \quad t(R[s]) = \text{fail}.\tag{2}

Limitations of Prior Work

ApproachLimitation
Manual audits (SWE-Bench+)Do not scale; must be repeated for every agent, model, and run
Automated auditors (ABA, BenchGuard)Audit the benchmark, not individual trials; verdicts rest on LLM judgment
Static test augmentation (UTBoost, SWE-ABS)Strengthen tests before any submission is observed; catch only anticipated failures
Dynamic differential testing (PatchDiff)Only 28.6% of divergent patches confirmable as incorrect; requires manual inspection

TESTJACK occupies the automated, dynamic quadrant: it audits individual trials, derives tests from the task prompt, uses ground truth only to validate expectations, and backs every verdict with an executable witness test.


Methodology

Per-Trial Auditing via Prompt-Grounded Differential Testing

TESTJACK implements auditing with two LLM agents (audit agent and reviewer agent) operating inside the benchmark's execution environment. The audit agent proceeds in five steps:

  1. Execution worlds: Build three repository copies — model world R[s]R[s], reference world R[G]R[G], and null world RR (without any patch, revealing which tests pass without a fix).

  2. Requirements checklist: Decompose PP into atomic, testable clauses instantiating Q(P)\mathcal{Q}(P), recording whether T0T_0 covers each clause and whether the trial implements it.

  3. Evidence collection: Examine the trial's diff along three axes:

    • Ghost branches: code paths never executed by T0T_0, identified via branch coverage
    • Reward hacking: detectors for hard-coded values, input special-casing, test-environment sniffing, harness tampering, stubs
    • General defects: specification violations, robustness/security issues, prioritized via mutation testing
  4. Issue-demonstrating tests: Write tests asserting the behavior each violated clause requires; tests must fail in the model world and pass in the reference world (oracle).

  5. Audit report: Assign verdicts (sound, incorrect despite passing, reward hack) with explanations of each defect, why it occurs, which clause requires correct behavior, and why T0T_0 fails to catch it.

The reviewer agent then re-judges these tests using PP as the sole authority (deliberately not using GG to infer requirements), labeling each as valid, questionable, or invalid. The trial is refuted if at least one valid test fails on it.

Lightweight Approach: Randomized Auditing with Witness Sharing

To reduce cost, TESTJACK offers a two-stage lightweight variant:

Stage 1 — Agentic task analysis: An agent reads PP, T0T_0, and GG, identifies requirements of Q(P)\mathcal{Q}(P) that T0T_0 does not adequately cover, and writes new tests targeting them. Tests failing on R[G]R[G] are refined using execution feedback or discarded after a refinement budget. This produces an augmented suite T1=T0∪TtaskT_1 = T_0 \cup T_{task}, built once per task.

Stage 2 — Randomized audit with cross-trial witness sharing:

  • Let TT denote the current test suite (initialized to T1T_1); trials failing any test in TT are refuted; remaining trials form the active set AA.
  • Loop: (i) sample trial ss uniformly from AA; (ii) run per-trial audit on ss; (iii) if witnesses WW are confirmed, add WW to TT, execute on every trial in AA, refuting all that fail; (iv) otherwise mark ss as clean and remove from AA.
  • Terminates when AA is empty or budget is exhausted.

Soundness of sharing: A witness belongs to the task rather than the trial on which it was found — the first two conditions of Eq.(2) do not involve ss at all, and the reviewer judges them against PP rather than the trial. Thus, deciding whether a witness refutes another trial reduces to the third condition.

The scheme is economical because generating/validating tests is expensive while replaying is cheap — the number of expensive audits approaches the number of distinct defects rather than defective trials.


Empirical Validation / Results

Experimental Setup

Benchmarks audited (Table 1):

Benchmark#TasksLanguagesTask origin#Passing trials
DeepSWE113TS, Go, Python, JS, RustAuthored from scratch2,848
SkillsBench*13VariousCommunity-authored55
SWE-Marathon‡2VariousExpert-authored28
SWE-bench Verified500PythonMined PRs805
SWE-bench Pro†731Python, Go, JS, TSMined commits751
Total1,3594,487

Agents: Claude Opus 4.6/4.7/4.8, GPT 5.5/5.6, Gemini 3.1 Pro. TESTJACK itself uses GPT-5.5 with high reasoning effort, capped at k=3k=3 generate-and-validate iterations per trial.

RQ1: Per-Trial Auditing Results

Table 2: Per-trial auditing results

Benchmark#All#Passing#RefutedRefuted (%)Corrected rate
DeepSWE6,2362,8481,24043.545.7% → 25.8%
SkillsBench85551018.264.7% → 52.9%
SWE-Marathon79281553.635.4% → 16.5%
SWE-bench Verified1,000805617.680.5% → 74.4%
SWE-bench Pro1,46275121929.251.4% → 36.4%
Total8,8624,4871,54534.450.6% → 33.2%

Key finding: Careful verifier design does not remove the problem — the from-scratch verifiers of DeepSWE and SWE-Marathon yield the two highest refutation rates (43.5% and 53.6%).

RQ2: Lightweight Approach Results

Table 3: Lightweight approach at 3 audit rounds per task

Benchmark#Per-trial#LightweightRecall (%)Prec. (%)Cost ($)‡Red. cost (%)
DeepSWE1,2401,82272.749.53,90484.9%
SkillsBench103180.025.88578.6%
SWE-Marathon151040.060.04853.7%
Total1,2651,86372.449.2†4,03784.6%

The lightweight approach recovers 72.4% of per-trial refutations at 15.4% of the cost. Limitations include: shared witnesses transfer well within a failure mode but rarely across modes; uniform sampling spends rounds on clean trials (hurting SWE-Marathon with only 28 passing trials).

RQ3: Comparison with SWE-ABS

Table 4: Comparison with SWE-ABS and ablation

BenchmarkMethod#RefutedRefuted (%)Recall (%)Precision (%)
DeepSWESWE-ABS2468.610.050.4
TESTJACK-Task Analysis34412.113.649.1
TESTJACK-Lightweight1,82264.072.749.5
SkillsBenchSWE-ABS23.610.050.0
TESTJACK-Task Analysis00.00.0-
TESTJACK-Lightweight3156.480.025.8
SWE-MarathonSWE-ABS00.00.0-
TESTJACK-Task Analysis13.66.7100.0
TESTJACK-Lightweight1035.740.060.0
TotalSWE-ABS2488.59.950.4
TESTJACK-Task Analysis34511.813.449.3
TESTJACK-Lightweight1,86363.672.449.2

TESTJACK-Lightweight achieves 7.3× the recall of SWE-ABS at comparable precision. On SWE-Marathon, SWE-ABS refutes no trials at all while TESTJACK refutes 10 of 28. Two design differences explain the gap: (1) SWE-ABS never observes submissions so must anticipate failures, while TESTJACK targets implementation-dependent defects; (2) SWE-ABS derives expectations from the gold patch rather than the prompt, penalizing implementation choices the prompt leaves open.

RQ4: Ablation Study

Task analysis alone refutes 345 trials (13.4% recall). Adding randomized auditing with witness sharing raises recall to 72.4% — a 5.4× increase — at essentially unchanged precision (49.3% vs. 49.2%). The stages are complementary: most refutations require evidence from a concrete submission.


Theoretical and Practical Implications

  1. Fundamental limitation of static evaluation: The results demonstrate that fixed test suites substantially overstate the capability of frontier coding agents. As LLMs become better at optimizing against fixed evaluators (reward hacking), those evaluators must themselves become more adaptive.

  2. Prompt-grounded differential testing as a paradigm: Separating what a test should check (from the prompt) from whether its expectations are correct (certified by the ground truth) provides a principled approach to dynamic evaluation that avoids both the test-oracle problem and overfitting to ground-truth implementation details.

  3. Cost-effective auditing at scale: The witness-sharing mechanism exploits the observation that trials for the same task often share failure modes, making expensive audits amortizable across many trials. This makes dynamic evaluation practical for large-scale benchmarks.

  4. Benchmark design implications: Even benchmarks built specifically to avoid weak tests (DeepSWE, SWE-Marathon) show the highest refutation rates, suggesting that the problem is not merely test quality but the inherent limitation of any static verifier against adaptive agents.

  5. Reference results for future work: The audit results can serve as references for future benchmark design and for recalibrating reported scores of existing benchmarks.


Conclusion

TESTJACK demonstrates that 34.4% of benchmark-passing trials from frontier coding agents violate task requirements, lowering overall resolution rates from 50.6% to 33.2% across five benchmarks. Even benchmarks designed specifically to avoid weak tests show the highest violation rates, revealing a fundamental limitation of current coding-agent evaluation.

The lightweight variant recovers 72.4% of refutations at 15.4% of the cost with 7.3× the recall of static augmentation methods, making dynamic evaluation practical at scale. Every refutation is backed by a replayable witness test that checks a prompt requirement, passes on the ground truth, and fails on the trial — providing executable evidence rather than LLM judgment alone.

Future directions implied by this work include: developing evaluators that evolve adaptively with the submissions they judge, improving the precision of lightweight auditing (particularly around underspecified prompts), and extending the framework to handle cross-failure-mode witness transfer and more efficient sampling strategies.

Related papers