Full text not available for this paper

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

Summary (Overview)

  • Core finding: A fixed LLM judge can make version-dependent errors when comparing agent upgrades, violating the implicit assumption that holding the judge constant makes judged differences reflect true differences in agent capability.

  • Multi-dataset evidence: The study analyzes 35 coding-agent submissions (20 version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories, finding that all available judges reject task-conditioned error invariance.

  • Decision-level consequences: 32 of 60 judge-by-pair units on SWE-bench show detectable differential comparison components, with 8 judge-only intervals declaring upgrades that execution-based intervals cannot establish—despite rank correlations of 0.71–0.79.

  • Transport fails: Rogan–Gladen calibration transported from a previous version amplifies comparison error (SWE-bench: from 3.8 to 19.5 points; ratio 5.20), because it divides differential error by the old version's Youden index.

  • Practical recommendation: Paired audits with power-tuned prediction-powered inference (PPI++) are preferred over judge-only release decisions or transported calibration; judges should be used for screening, not final release decisions.

Introduction and Theoretical Foundation

Background and Motivation

Software teams increasingly decide whether to ship a new agent version by comparing it with the current one under an automatic evaluator—most often another LLM acting as a judge. This practice rests on a rarely-stated assumption: if the judge is held fixed, a difference in judged success between two versions reflects a difference in their actual success.

The paper builds on classical measurement-error theory:

  • Non-differential misclassification attenuates a comparison toward zero
  • Differential misclassification can bias it in either direction (Bross, 1954; Copeland et al., 1977)

Prior work (Dorner et al., 2025; Fiedler, 2026) showed that model-dependent judge bias can reverse rankings and that shared calibration amplifies bias. However, what remained unclear was how much this matters for the actual decision: comparing a new version with its predecessor on execution-verifiable benchmarks when versions are close in quality.

Key Theoretical Framework

The paper distinguishes two notions of non-differential error:

  1. Conditionally non-differential: For every task tt and label hh, P(Z=1∣H=h,task t)P(Z = 1 | H = h, \text{task } t) is the same whichever version produced the trajectory
  2. Marginally non-differential: TPRo=TPRnTPR_o = TPR_n and FPRo=FPRnFPR_o = FPR_n

Neither implies the other—a crucial distinction for interpreting judge behavior.

Methodology

Datasets

DatasetReference StandardScale
SWE-bench Verified (primary)Execution-based test resolution35 agents, 250 tasks, 20 version pairs
τ-benchDatabase-state environment reward2 agents, 155 tasks, 4 trials each
AgentRewardBenchExpert annotations4 agents, 300 tasks, 1,106 trajectories

Judges

Four judge models from three providers were queried:

  • GPT-5 mini and GPT-6 Sol (OpenAI)
  • Gemini 3.5 Flash (Google)
  • Claude Opus 5.5 (Anthropic)

An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned cells); the fourth (Claude) is reported descriptively on its matched partial sample.

Key Propositions

Proposition 1 (Attenuation and reversal):

e=δ−(1−Jo)DHe = \delta - (1 - J_o) D_H

where δ=ΔTPR⋅pn+ΔFPR⋅(1−pn)\delta = \Delta TPR \cdot p_n + \Delta FPR \cdot (1 - p_n), ΔTPR=TPRn−TPRo\Delta TPR = TPR_n - TPR_o, and ΔFPR=FPRn−FPRo\Delta FPR = FPR_n - FPR_o. Close version pairs have small DHD_H, so their comparison error is dominated by δ\delta.

Proposition 2 (Transported calibration):

eT=p^n−pn=δJoe_T = \hat{p}_n - p_n = \frac{\delta}{J_o}

Transport removes attenuation but divides the differential component by Jo<1J_o < 1—making close comparisons worse when ∣DH∣|D_H| is small relative to ∣δ∣|\delta|.

Proposition 3 (Paired-difference efficiency):

τD−1=(τL−1)1−r1−rε\tau_D - 1 = (\tau_L - 1) \frac{1 - r}{1 - r_\varepsilon}

where r=corr(Ho,Hn)r = \text{corr}(H_o, H_n) is the task-level correlation of the two versions' outcomes and rε=corr(εo,εn)r_\varepsilon = \text{corr}(\varepsilon_o, \varepsilon_n) that of the judge's errors.

Proposition 4 (Attenuation is set by tasks that differ):

DJ=N+J+−N−J−N=JˉDH+Covt(Jt,Dt)D_J = \frac{N_+ J_+ - N_- J_-}{N} = \bar{J} D_H + \text{Cov}_t(J_t, D_t)

When J+=J−=JΔJ_+ = J_- = J_\Delta, a conditionally non-differential judge attenuates by its discriminability on discordant tasks.

Hypotheses

  • H1: Judge verdicts depend on the agent after conditioning on task and reference label
  • H2: Biased version comparisons—differential component differs from zero more often than chance
  • H3: Transport fails—leaves more than half of naive comparison error
  • H4: Paired audits have valid coverage where judge-only intervals do not
  • H5: Self-report mechanism—adding agent's final message raises false acceptance
  • H6: Capability gradient—false-positive contrast rises with agent capability
  • H7: Solvability gradient—acceptance of failed patches increases with other agents' success

Empirical Validation / Results

AgentRewardBench (Replication)

  • H1: 5 of 14 LLM evaluators reject non-differential error after Holm adjustment (fails half-of-judges criterion)
  • Task-conditioned contrasts: All 15 evaluators accept Llama 3.3's failed trajectories less often (mean −6.7 points) and Claude 3.7 Sonnet's failures more often (mean +5.0 points)
  • Transport ratio: 1.55 (95% CI 1.11–2.11)—worse than uncorrected judge in 58.3% of comparisons
  • Paired efficiency: 1.30 vs. 1.64 for levels (median paired gain is 45% of level gain)

SWE-bench Verified (Primary)

Differential error (H1): All three available judges reject the null (p = 0.0001 per judge). Median false-positive rates: 65.7% (Gemini), 66.0% (GPT-5 mini), 40.0% (GPT-6 Sol).

Capability gradient (H6): Mean Spearman correlation of resolve rate with false-positive contrast: +0.819 (p = 0.0001). Youden contrast: −0.747 (p = 0.0001).

Solvability gradient (H7): Fails in the opposite direction—false acceptance decreases by 38–48 points as other agents' success share rises from zero to one.

Version comparisons (H2):

JudgeMedian FPR / FNRNonzero ede_d after FDRRelease-decision disagreementsMean absolute eeMean absolute transported error
Gemini 3.5 Flash65.7% / 9.6%12/204/203.3 points20.3 points
GPT-5 mini66.0% / 13.1%9/204/202.9 points21.8 points
GPT-6 Sol40.0% / 22.2%11/203/205.1 points16.5 points

Transport (H3): Mean absolute error rises from 3.8 to 19.5 points (R = 5.20). 24.6% of bootstrap draws undefined due to nonpositive Youden indices. All 200 simulated non-differential null ratios are exceeded (p = 0.005).

Audits (H4): PPI++ narrows width by only ~5% at 80 tasks (14.2 vs. 14.9 points). Judge-only intervals cover the reference in only 80.0% of units.

τ-bench

  • All four judges reject H1 (cluster-robust p ≤ 0.0015)
  • GPT-6 Sol reverses a 9-point reference gap: judges GPT-4o 17.8 points better (CI 11.8–24.1) when the reward shows it 9.0 points worse
  • Mechanism identified: GPT-6 Sol penalizes tool-and-message turns that violate policy but are ignored by the environment reward (31.7% of Claude rejections cite this rule vs. 0% for GPT-4o)

Randomized Self-Report Mechanism (H5)

The predicted false-acceptance increase was not observed:

JudgeEffect (points)95% CIAdjusted p
Gemini 3.5 Flash−1.47(−4.16, +0.99)1.0
GPT-5 mini−0.49(−2.84, +1.67)1.0
GPT-6 Sol+0.61(−1.83, +3.51)1.0

Post-Submission OpenHands Validation

Eight OpenHands configurations on the same 250 issues reproduce the capability/error association:

JudgeResolve rate vs. FPR contrastResolve rate vs. Youden contrast
Gemini 3.5 Flash+0.905−0.976
GPT-5 mini+0.976−0.952
GPT-6 Sol+0.952−0.881
Mean+0.944 (exact p = 0.000099)−0.937 (p = 0.000397)

Theoretical and Practical Implications

Theoretical Contributions

  1. Exact decomposition of transport failure: Proposition 2 shows why shared calibration amplifies differential error by 1/Jo1/J_o, connecting classical measurement-error theory to LLM-judge evaluation.

  2. Paired-difference efficiency bound: Proposition 3 quantifies how much a judge can help for comparisons vs. levels, governed by (1−r)/(1−rε)(1-r)/(1-r_\varepsilon).

  3. Discordant-task attenuation: Proposition 4 shows attenuation depends on judge discriminability on tasks where versions differ, not average discriminability.

Practical Guidance

  1. Treat judged differences as screening, not evidence
  2. Audit the comparison, not the level—use PPI++ with randomly sampled reference labels
  3. Size audits from the comparison's own correlation—close versions need large audits; a few hundred tasks cannot resolve a 5-point difference
  4. Do not transport calibration unless invariance is checked on the new version
  5. Keep anchors for judge monitoring, not validity checks
  6. Report class-conditional judge error by version
  7. Test rather than assume self-report effects

Audit Sizing Formula

n≳[1N+Δ2(z1−α/2+z1−β)2S2(1−ρD2)]−1n \gtrsim \left[ \frac{1}{N} + \frac{\Delta^2}{(z_{1-\alpha/2} + z_{1-\beta})^2 S^2 (1 - \rho_D^2)} \right]^{-1}

For a 5-point difference with 80% power in a 300-task pool: ~217 labels without a judge, ~202 with a judge (ρD=0.46\rho_D = 0.46)—a saving of only 7%.

Conclusion

Using a frozen LLM judge does not make comparisons between agent versions valid. Key takeaways:

  1. Differential error is pervasive: All available judges show task-conditioned, agent-dependent error across all datasets.

  2. Release decisions can be wrong: Eight judge-only SWE-bench upgrades were not established by execution-based intervals; τ-bench showed a confident reversal.

  3. Transport is fragile: Old-version calibration amplifies error by factors of 1.55–5.20 across datasets.

  4. Paired audits are valid but label-efficient only modestly: PPI++ saves ~5–13% in interval width for close versions.

  5. No universal mechanism explains agent-dependent error: The solvability gradient reverses sign across domains; the self-report intervention fails its prespecified test.

Final recommendation: State the reference standard, report version- and task-conditioned error, use judges for screening, and base release decisions on paired audits of current outputs—not on frozen-anchor scores or transported calibration alone. Independent human patch review remains pending.

Related papers