# Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

> A fixed LLM judge produces version-dependent errors, invalidating agent comparisons and transported calibration, so release decisions require paired audits, not judge-only scores.

- **Source:** [arXiv](https://arxiv.org/abs/2609.34198)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/X6LwmV
- **Whiteboard:** https://picx.dev/p/X6LwmV/image

## Summary

# Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

## Summary (Overview)

- **Core finding**: A fixed LLM judge can make version-dependent errors when comparing agent upgrades, violating the implicit assumption that holding the judge constant makes judged differences reflect true differences in agent capability.

- **Multi-dataset evidence**: The study analyzes 35 coding-agent submissions (20 version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories, finding that all available judges reject task-conditioned error invariance.

- **Decision-level consequences**: 32 of 60 judge-by-pair units on SWE-bench show detectable differential comparison components, with 8 judge-only intervals declaring upgrades that execution-based intervals cannot establish—despite rank correlations of 0.71–0.79.

- **Transport fails**: Rogan–Gladen calibration transported from a previous version amplifies comparison error (SWE-bench: from 3.8 to 19.5 points; ratio 5.20), because it divides differential error by the old version's Youden index.

- **Practical recommendation**: Paired audits with power-tuned prediction-powered inference (PPI++) are preferred over judge-only release decisions or transported calibration; judges should be used for screening, not final release decisions.

## Introduction and Theoretical Foundation

### Background and Motivation

Software teams increasingly decide whether to ship a new agent version by comparing it with the current one under an automatic evaluator—most often another LLM acting as a judge. This practice rests on a rarely-stated assumption: **if the judge is held fixed, a difference in judged success between two versions reflects a difference in their actual success**.

The paper builds on classical measurement-error theory:

- **Non-differential misclassification** attenuates a comparison toward zero
- **Differential misclassification** can bias it in either direction (Bross, 1954; Copeland et al., 1977)

Prior work (Dorner et al., 2025; Fiedler, 2026) showed that model-dependent judge bias can reverse rankings and that shared calibration amplifies bias. However, what remained unclear was **how much this matters for the actual decision**: comparing a new version with its predecessor on execution-verifiable benchmarks when versions are close in quality.

### Key Theoretical Framework

The paper distinguishes two notions of non-differential error:

1. **Conditionally non-differential**: For every task $t$ and label $h$, $P(Z = 1 | H = h, \text{task } t)$ is the same whichever version produced the trajectory
2. **Marginally non-differential**: $TPR_o = TPR_n$ and $FPR_o = FPR_n$

Neither implies the other—a crucial distinction for interpreting judge behavior.

## Methodology

### Datasets

| Dataset | Reference Standard | Scale |
|---------|-------------------|-------|
| SWE-bench Verified (primary) | Execution-based test resolution | 35 agents, 250 tasks, 20 version pairs |
| τ-bench | Database-state environment reward | 2 agents, 155 tasks, 4 trials each |
| AgentRewardBench | Expert annotations | 4 agents, 300 tasks, 1,106 trajectories |

### Judges

Four judge models from three providers were queried:
- GPT-5 mini and GPT-6 Sol (OpenAI)
- Gemini 3.5 Flash (Google)
- Claude Opus 5.5 (Anthropic)

An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned cells); the fourth (Claude) is reported descriptively on its matched partial sample.

### Key Propositions

**Proposition 1 (Attenuation and reversal)**:

$$e = \delta - (1 - J_o) D_H$$

where $\delta = \Delta TPR \cdot p_n + \Delta FPR \cdot (1 - p_n)$, $\Delta TPR = TPR_n - TPR_o$, and $\Delta FPR = FPR_n - FPR_o$. Close version pairs have small $D_H$, so their comparison error is dominated by $\delta$.

**Proposition 2 (Transported calibration)**:

$$e_T = \hat{p}_n - p_n = \frac{\delta}{J_o}$$

Transport removes attenuation but divides the differential component by $J_o < 1$—making close comparisons worse when $|D_H|$ is small relative to $|\delta|$.

**Proposition 3 (Paired-difference efficiency)**:

$$\tau_D - 1 = (\tau_L - 1) \frac{1 - r}{1 - r_\varepsilon}$$

where $r = \text{corr}(H_o, H_n)$ is the task-level correlation of the two versions' outcomes and $r_\varepsilon = \text{corr}(\varepsilon_o, \varepsilon_n)$ that of the judge's errors.

**Proposition 4 (Attenuation is set by tasks that differ)**:

$$D_J = \frac{N_+ J_+ - N_- J_-}{N} = \bar{J} D_H + \text{Cov}_t(J_t, D_t)$$

When $J_+ = J_- = J_\Delta$, a conditionally non-differential judge attenuates by its discriminability on discordant tasks.

### Hypotheses

- **H1**: Judge verdicts depend on the agent after conditioning on task and reference label
- **H2**: Biased version comparisons—differential component differs from zero more often than chance
- **H3**: Transport fails—leaves more than half of naive comparison error
- **H4**: Paired audits have valid coverage where judge-only intervals do not
- **H5**: Self-report mechanism—adding agent's final message raises false acceptance
- **H6**: Capability gradient—false-positive contrast rises with agent capability
- **H7**: Solvability gradient—acceptance of failed patches increases with other agents' success

## Empirical Validation / Results

### AgentRewardBench (Replication)

- **H1**: 5 of 14 LLM evaluators reject non-differential error after Holm adjustment (fails half-of-judges criterion)
- **Task-conditioned contrasts**: All 15 evaluators accept Llama 3.3's failed trajectories less often (mean −6.7 points) and Claude 3.7 Sonnet's failures more often (mean +5.0 points)
- **Transport ratio**: 1.55 (95% CI 1.11–2.11)—worse than uncorrected judge in 58.3% of comparisons
- **Paired efficiency**: 1.30 vs. 1.64 for levels (median paired gain is 45% of level gain)

### SWE-bench Verified (Primary)

**Differential error (H1)**: All three available judges reject the null (p = 0.0001 per judge). Median false-positive rates: 65.7% (Gemini), 66.0% (GPT-5 mini), 40.0% (GPT-6 Sol).

**Capability gradient (H6)**: Mean Spearman correlation of resolve rate with false-positive contrast: **+0.819** (p = 0.0001). Youden contrast: **−0.747** (p = 0.0001).

**Solvability gradient (H7)**: **Fails in the opposite direction**—false acceptance decreases by 38–48 points as other agents' success share rises from zero to one.

**Version comparisons (H2)**:

| Judge | Median FPR / FNR | Nonzero $e_d$ after FDR | Release-decision disagreements | Mean absolute $e$ | Mean absolute transported error |
|-------|------------------|------------------------|-------------------------------|-------------------|--------------------------------|
| Gemini 3.5 Flash | 65.7% / 9.6% | 12/20 | 4/20 | 3.3 points | 20.3 points |
| GPT-5 mini | 66.0% / 13.1% | 9/20 | 4/20 | 2.9 points | 21.8 points |
| GPT-6 Sol | 40.0% / 22.2% | 11/20 | 3/20 | 5.1 points | 16.5 points |

**Transport (H3)**: Mean absolute error rises from 3.8 to 19.5 points (R = 5.20). 24.6% of bootstrap draws undefined due to nonpositive Youden indices. All 200 simulated non-differential null ratios are exceeded (p = 0.005).

**Audits (H4)**: PPI++ narrows width by only ~5% at 80 tasks (14.2 vs. 14.9 points). Judge-only intervals cover the reference in only 80.0% of units.

### τ-bench

- **All four judges reject H1** (cluster-robust p ≤ 0.0015)
- **GPT-6 Sol reverses a 9-point reference gap**: judges GPT-4o 17.8 points better (CI 11.8–24.1) when the reward shows it 9.0 points worse
- **Mechanism identified**: GPT-6 Sol penalizes tool-and-message turns that violate policy but are ignored by the environment reward (31.7% of Claude rejections cite this rule vs. 0% for GPT-4o)

### Randomized Self-Report Mechanism (H5)

The predicted false-acceptance increase was **not observed**:

| Judge | Effect (points) | 95% CI | Adjusted p |
|-------|----------------|--------|------------|
| Gemini 3.5 Flash | −1.47 | (−4.16, +0.99) | 1.0 |
| GPT-5 mini | −0.49 | (−2.84, +1.67) | 1.0 |
| GPT-6 Sol | +0.61 | (−1.83, +3.51) | 1.0 |

### Post-Submission OpenHands Validation

Eight OpenHands configurations on the same 250 issues reproduce the capability/error association:

| Judge | Resolve rate vs. FPR contrast | Resolve rate vs. Youden contrast |
|-------|------------------------------|----------------------------------|
| Gemini 3.5 Flash | +0.905 | −0.976 |
| GPT-5 mini | +0.976 | −0.952 |
| GPT-6 Sol | +0.952 | −0.881 |
| **Mean** | **+0.944** (exact p = 0.000099) | **−0.937** (p = 0.000397) |

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Exact decomposition of transport failure**: Proposition 2 shows why shared calibration amplifies differential error by $1/J_o$, connecting classical measurement-error theory to LLM-judge evaluation.

2. **Paired-difference efficiency bound**: Proposition 3 quantifies how much a judge can help for comparisons vs. levels, governed by $(1-r)/(1-r_\varepsilon)$.

3. **Discordant-task attenuation**: Proposition 4 shows attenuation depends on judge discriminability on tasks where versions differ, not average discriminability.

### Practical Guidance

1. **Treat judged differences as screening, not evidence**
2. **Audit the comparison, not the level**—use PPI++ with randomly sampled reference labels
3. **Size audits from the comparison's own correlation**—close versions need large audits; a few hundred tasks cannot resolve a 5-point difference
4. **Do not transport calibration** unless invariance is checked on the new version
5. **Keep anchors for judge monitoring, not validity checks**
6. **Report class-conditional judge error by version**
7. **Test rather than assume self-report effects**

### Audit Sizing Formula

$$n \gtrsim \left[ \frac{1}{N} + \frac{\Delta^2}{(z_{1-\alpha/2} + z_{1-\beta})^2 S^2 (1 - \rho_D^2)} \right]^{-1}$$

For a 5-point difference with 80% power in a 300-task pool: ~217 labels without a judge, ~202 with a judge ($\rho_D = 0.46$)—a saving of only 7%.

## Conclusion

Using a frozen LLM judge does **not** make comparisons between agent versions valid. Key takeaways:

1. **Differential error is pervasive**: All available judges show task-conditioned, agent-dependent error across all datasets.

2. **Release decisions can be wrong**: Eight judge-only SWE-bench upgrades were not established by execution-based intervals; τ-bench showed a confident reversal.

3. **Transport is fragile**: Old-version calibration amplifies error by factors of 1.55–5.20 across datasets.

4. **Paired audits are valid but label-efficient only modestly**: PPI++ saves ~5–13% in interval width for close versions.

5. **No universal mechanism explains agent-dependent error**: The solvability gradient reverses sign across domains; the self-report intervention fails its prespecified test.

**Final recommendation**: State the reference standard, report version- and task-conditioned error, use judges for screening, and base release decisions on paired audits of current outputs—not on frozen-anchor scores or transported calibration alone. Independent human patch review remains pending.

---

_Markdown view of https://picx.dev/p/X6LwmV, served by PicX — AI-generated visual whiteboard summaries of research papers._
