Full text not available for this paper
Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
Summary (Overview)
-
Core finding: A fixed LLM judge can make version-dependent errors when comparing agent upgrades, violating the implicit assumption that holding the judge constant makes judged differences reflect true differences in agent capability.
-
Multi-dataset evidence: The study analyzes 35 coding-agent submissions (20 version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories, finding that all available judges reject task-conditioned error invariance.
-
Decision-level consequences: 32 of 60 judge-by-pair units on SWE-bench show detectable differential comparison components, with 8 judge-only intervals declaring upgrades that execution-based intervals cannot establish—despite rank correlations of 0.71–0.79.
-
Transport fails: Rogan–Gladen calibration transported from a previous version amplifies comparison error (SWE-bench: from 3.8 to 19.5 points; ratio 5.20), because it divides differential error by the old version's Youden index.
-
Practical recommendation: Paired audits with power-tuned prediction-powered inference (PPI++) are preferred over judge-only release decisions or transported calibration; judges should be used for screening, not final release decisions.
Introduction and Theoretical Foundation
Background and Motivation
Software teams increasingly decide whether to ship a new agent version by comparing it with the current one under an automatic evaluator—most often another LLM acting as a judge. This practice rests on a rarely-stated assumption: if the judge is held fixed, a difference in judged success between two versions reflects a difference in their actual success.
The paper builds on classical measurement-error theory:
- Non-differential misclassification attenuates a comparison toward zero
- Differential misclassification can bias it in either direction (Bross, 1954; Copeland et al., 1977)
Prior work (Dorner et al., 2025; Fiedler, 2026) showed that model-dependent judge bias can reverse rankings and that shared calibration amplifies bias. However, what remained unclear was how much this matters for the actual decision: comparing a new version with its predecessor on execution-verifiable benchmarks when versions are close in quality.
Key Theoretical Framework
The paper distinguishes two notions of non-differential error:
- Conditionally non-differential: For every task and label , is the same whichever version produced the trajectory
- Marginally non-differential: and
Neither implies the other—a crucial distinction for interpreting judge behavior.
Methodology
Datasets
| Dataset | Reference Standard | Scale |
|---|---|---|
| SWE-bench Verified (primary) | Execution-based test resolution | 35 agents, 250 tasks, 20 version pairs |
| τ-bench | Database-state environment reward | 2 agents, 155 tasks, 4 trials each |
| AgentRewardBench | Expert annotations | 4 agents, 300 tasks, 1,106 trajectories |
Judges
Four judge models from three providers were queried:
- GPT-5 mini and GPT-6 Sol (OpenAI)
- Gemini 3.5 Flash (Google)
- Claude Opus 5.5 (Anthropic)
An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned cells); the fourth (Claude) is reported descriptively on its matched partial sample.
Key Propositions
Proposition 1 (Attenuation and reversal):
where , , and . Close version pairs have small , so their comparison error is dominated by .
Proposition 2 (Transported calibration):
Transport removes attenuation but divides the differential component by —making close comparisons worse when is small relative to .
Proposition 3 (Paired-difference efficiency):
where is the task-level correlation of the two versions' outcomes and that of the judge's errors.
Proposition 4 (Attenuation is set by tasks that differ):
When , a conditionally non-differential judge attenuates by its discriminability on discordant tasks.
Hypotheses
- H1: Judge verdicts depend on the agent after conditioning on task and reference label
- H2: Biased version comparisons—differential component differs from zero more often than chance
- H3: Transport fails—leaves more than half of naive comparison error
- H4: Paired audits have valid coverage where judge-only intervals do not
- H5: Self-report mechanism—adding agent's final message raises false acceptance
- H6: Capability gradient—false-positive contrast rises with agent capability
- H7: Solvability gradient—acceptance of failed patches increases with other agents' success
Empirical Validation / Results
AgentRewardBench (Replication)
- H1: 5 of 14 LLM evaluators reject non-differential error after Holm adjustment (fails half-of-judges criterion)
- Task-conditioned contrasts: All 15 evaluators accept Llama 3.3's failed trajectories less often (mean −6.7 points) and Claude 3.7 Sonnet's failures more often (mean +5.0 points)
- Transport ratio: 1.55 (95% CI 1.11–2.11)—worse than uncorrected judge in 58.3% of comparisons
- Paired efficiency: 1.30 vs. 1.64 for levels (median paired gain is 45% of level gain)
SWE-bench Verified (Primary)
Differential error (H1): All three available judges reject the null (p = 0.0001 per judge). Median false-positive rates: 65.7% (Gemini), 66.0% (GPT-5 mini), 40.0% (GPT-6 Sol).
Capability gradient (H6): Mean Spearman correlation of resolve rate with false-positive contrast: +0.819 (p = 0.0001). Youden contrast: −0.747 (p = 0.0001).
Solvability gradient (H7): Fails in the opposite direction—false acceptance decreases by 38–48 points as other agents' success share rises from zero to one.
Version comparisons (H2):
| Judge | Median FPR / FNR | Nonzero after FDR | Release-decision disagreements | Mean absolute | Mean absolute transported error |
|---|---|---|---|---|---|
| Gemini 3.5 Flash | 65.7% / 9.6% | 12/20 | 4/20 | 3.3 points | 20.3 points |
| GPT-5 mini | 66.0% / 13.1% | 9/20 | 4/20 | 2.9 points | 21.8 points |
| GPT-6 Sol | 40.0% / 22.2% | 11/20 | 3/20 | 5.1 points | 16.5 points |
Transport (H3): Mean absolute error rises from 3.8 to 19.5 points (R = 5.20). 24.6% of bootstrap draws undefined due to nonpositive Youden indices. All 200 simulated non-differential null ratios are exceeded (p = 0.005).
Audits (H4): PPI++ narrows width by only ~5% at 80 tasks (14.2 vs. 14.9 points). Judge-only intervals cover the reference in only 80.0% of units.
τ-bench
- All four judges reject H1 (cluster-robust p ≤ 0.0015)
- GPT-6 Sol reverses a 9-point reference gap: judges GPT-4o 17.8 points better (CI 11.8–24.1) when the reward shows it 9.0 points worse
- Mechanism identified: GPT-6 Sol penalizes tool-and-message turns that violate policy but are ignored by the environment reward (31.7% of Claude rejections cite this rule vs. 0% for GPT-4o)
Randomized Self-Report Mechanism (H5)
The predicted false-acceptance increase was not observed:
| Judge | Effect (points) | 95% CI | Adjusted p |
|---|---|---|---|
| Gemini 3.5 Flash | −1.47 | (−4.16, +0.99) | 1.0 |
| GPT-5 mini | −0.49 | (−2.84, +1.67) | 1.0 |
| GPT-6 Sol | +0.61 | (−1.83, +3.51) | 1.0 |
Post-Submission OpenHands Validation
Eight OpenHands configurations on the same 250 issues reproduce the capability/error association:
| Judge | Resolve rate vs. FPR contrast | Resolve rate vs. Youden contrast |
|---|---|---|
| Gemini 3.5 Flash | +0.905 | −0.976 |
| GPT-5 mini | +0.976 | −0.952 |
| GPT-6 Sol | +0.952 | −0.881 |
| Mean | +0.944 (exact p = 0.000099) | −0.937 (p = 0.000397) |
Theoretical and Practical Implications
Theoretical Contributions
-
Exact decomposition of transport failure: Proposition 2 shows why shared calibration amplifies differential error by , connecting classical measurement-error theory to LLM-judge evaluation.
-
Paired-difference efficiency bound: Proposition 3 quantifies how much a judge can help for comparisons vs. levels, governed by .
-
Discordant-task attenuation: Proposition 4 shows attenuation depends on judge discriminability on tasks where versions differ, not average discriminability.
Practical Guidance
- Treat judged differences as screening, not evidence
- Audit the comparison, not the level—use PPI++ with randomly sampled reference labels
- Size audits from the comparison's own correlation—close versions need large audits; a few hundred tasks cannot resolve a 5-point difference
- Do not transport calibration unless invariance is checked on the new version
- Keep anchors for judge monitoring, not validity checks
- Report class-conditional judge error by version
- Test rather than assume self-report effects
Audit Sizing Formula
For a 5-point difference with 80% power in a 300-task pool: ~217 labels without a judge, ~202 with a judge ()—a saving of only 7%.
Conclusion
Using a frozen LLM judge does not make comparisons between agent versions valid. Key takeaways:
-
Differential error is pervasive: All available judges show task-conditioned, agent-dependent error across all datasets.
-
Release decisions can be wrong: Eight judge-only SWE-bench upgrades were not established by execution-based intervals; τ-bench showed a confident reversal.
-
Transport is fragile: Old-version calibration amplifies error by factors of 1.55–5.20 across datasets.
-
Paired audits are valid but label-efficient only modestly: PPI++ saves ~5–13% in interval width for close versions.
-
No universal mechanism explains agent-dependent error: The solvability gradient reverses sign across domains; the self-report intervention fails its prespecified test.
Final recommendation: State the reference standard, report version- and task-conditioned error, use judges for screening, and base release decisions on paired audits of current outputs—not on frozen-anchor scores or transported calibration alone. Independent human patch review remains pending.
Related papers
- Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
Identical runs of the same AI coding agent vary more than differences between agent-model pairings, so best-of-three with compliance checking beats single-run benchmarking.
- How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Under matched conditions, a minimal-harness coding agent matches or outperforms state-of-the-art MLE harnesses, with performance driven by the LLM backbone and execution environment, not scaffolding.
- From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining
Search scaling in autonomous research is capability-dependent: deeper search narrows cross-model gaps, but early trajectory state and parallel exploration critically shape final performance.