# What pass@k Cannot Measure: Evaluating Diversity and Capability Retention after Post-Training

> pass@k cannot detect diversity loss: GRPO and RFT move entropy and answer diversity in opposite directions with zero seed overlap, yet pass@8 and pass@32 show no consistent winner.

- **Source:** [arXiv](https://arxiv.org/abs/2610.07405)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/PlBn2A
- **Whiteboard:** https://picx.dev/p/PlBn2A/image

## Summary

## Summary (Overview)

- **Core finding**: pass@k, the standard evaluation metric for post-training quality, is structurally blind to how probability mass is distributed across distinct outputs. It measures the quantity of correct-answer mass, not its diversity or organization.
- **Empirical demonstration**: Training Qwen2.5-1.5B-Instruct on GSM8K math problems with GRPO (Group Relative Policy Optimization) and RFT (Rejection-Sampling Fine-tuning) moves three complementary diversity measures (token entropy, answer entropy, unique answers per prompt)in opposite directions with zero overlap across three seeds per arm
- **Metric failure**: pass@8 and pass@32 show no consistent winner between GRPO and RFT, despite large, seed-robust diversity differences. On a hard MATH-500 subset, arms separate only at low k and converge as k grows
- **Missing control problem**: Compared against the starting checkpoint, no trained arm significantly improves hard-problem coverage. RFT is significantly worse than baseline, while GRPO is statistically indistinguishable from it, meaning GRPO's pass@1 edge over RFT reflects a smaller loss, not a capability gain
- **Key implication**: An improved pass@k cannot certify that post-training preserved, rather than narrowed, what the starting checkpoint could do. The metric has a construct-validity gap that is structural, not an implementation artifact

## Introduction and Theoretical Foundation

The paper addresses a fundamental question in post-training evaluation: how do we know whether reinforcement learning with verifiable rewards (RLVR) converts a broad-coverage base model into one that reliably solves tasks without secretly narrowing its capabilities?

**The pass@k metric**: The field's default protocol uses pass@k [Chen et al.,, 2021],the probability that at least one of k independently sampled attempts at a problem is correct, computed with the unbiased estimator:

$$1 - \frac{\binom{n-c}{k}}{\binom{n}{k}}$$

over $n \geq k$ samples with c correct. pass@1 is treated as capability;pass@k at larger k is implicitly treated as a check on whether that capability cost the model's solution diversity.



**Structural construct-validity gap**: At the population level, for a problem with true per-sample correctness probability $p$,:

$$\text{pass@k}(p) = 1 - (1-p)^k$$

a function of p alone. Two models with the same p have the same pass@k at every k, regardless of how the remaining probability mass is arranged across distinct correct solutions, incorrect ones, or reasoning paths. The paper states:

> "pass@k is a statement about how much correct mass exists, not about how it is distributed. A post-training method that redistributes mass among already-correct outputs, or among incorrect ones, without changing p much, is therefore invisible to it by construction rather than by an implementation gap."

**Related work and disagreements**: The literature disagrees on whether RL narrows or broadens diversity:
- **Yue et al.**[2025]: RLVR-trained models beat base models at small k but are overtaken once k reaches tens to hundreds, across six RL algorithms
- **Nicolicioiu et al.**[2026]: On-policy self-distillation raises pass@1 while flattening pass@16,through a mechanism distinct from RL mode-seeking
- **Kim et al.**[2026]: GRPO increases token entropy on two 7–8B models, while self-distillation decreases it—the reverse of what this paper reports at smaller scale

This paper reframes the contribution: rather than claiming "RL narrows diversity"(already disputed),it claims the protocol itself—pass@k at the k values the field reports—is not a valid standalone proxy for diversity,independent of which direction any method moves it.



## Methodology

**Model and data**:
- **Base model**: Qwen2.5-1.5B-Instruct (pre-post-training baseline, already instruction-tuned)
- **Training data**: Frozen pool of 1,500 GSM8K problems [Cobbe et al.,, 2021], filtered to keep problems where 1–7 of 8 sampled rollouts land correct
- **Evaluation**: In-domain on held-out 100 GSM8K problems(32 samples/problem); out-of-domain on MATH-500 [Lightman et al,, 2023] (16 samples/problem, never trained on)

**Training arms**:
1. **GRPO** [Shao et al,, 2024]: group-relative advantage over 8 rollouts/prompt, clipped policy-gradient update, no KL penalty
2. **RFT** (rejection-sampling fine-tuning): 8 rollouts/prompt, keep only the shortest verifier-passed one, fine-tuned with plain cross-entropy
3. **RFT + KL brake**: adds Kullback-Leibler penalty toward the frozen starting policy
4. **Context arms** (single seed each): GRPO with length penalties(λ = 0.2, 0.5)and on-policy self-distillation(OPD)

All generation uses temperature 1.0, top-p 1.0, no top-k truncation, 1,024-token budget. Each arm runs at three seeds(0, 1, 2),187 steps each

**Metrics**:
- **Token-level entropy**, **answer-level entropy**, **unique answers per prompt**: computed on the policy model's own tokenizer output
- **pass@1/8/32**: each seed's own value carries 95% bootstrap CI;3-seed summaries combine intervals with DerSimonian–Laird random-effects meta-analysis
- **Distinct-3** [Li et al,, 2016]: fraction of unique word-trigrams pooled across verifier-correct completions only, plus a length-truncated variant
- **$D_{\mathrm{incorrect}}$** : fraction of unique incorrect answers among incorrect completions only, plus a count-controlled variant(subsampling exactly 4 incorrect samples per problem)
- **Hard-problem evaluation**: MATH-500 restricted to hardest levels(4–5,100 problems),pass@k for k ∈ {1,2,4,8,16},generated against each final checkpoint and the pre-intervention checkpoint

## Empirical Validation / Results

**1. Diversity moves in opposite, seed-robust directions**

**Table 1: GSM8K diversity metrics, change from first to last checkpoint, mean ± half-range across 3 seeds**

| Arm | Token entropy Δ | Unique answers/problem Δ | Answer entropy Δ |
|---|---|---|---|
| RFT (shortest-correct) | +0.028 ± 0.003 | +0.56 ± 0.06 | +0.091 ± 0.017 |
| RFT + KL brake | +0.032 ± 0.001 | +0.52 ± 0.11 | +0.082 ±  ​​0.009 |
| GRPO | -0.055 ± 0.004 | -1.30 ± 0.20 | -0.164 ±  ​​0.018 |
| *Single seed each;context only*: | | | |
| GRPO + length penalty (λ=0.2) | -0.055 | -1.26 | -0.163 |
| GRPO + length penalty (λ=0.5) | -0.050 | -1.20 | -0.150 |
| On-policy self-distillation (OPD) | -0.003 | -0.34 | -0.035 |

Both RFT arms increase all three measures in every seed;GRPO decreases all three in every seed,with no cross-arm overlap. Notably,GRPO narrows diversity despite being the arm least resembling self-distillation,contradicting a mode-seeking account.



**2. pass@1 separates arms cleanly;pass@8/32 do not**

**Table 2: GSM8K pass@k, final checkpoint(3-seed DerSimonian–Laird pooled mean,95% CI)**

| Arm | pass@1 | pass@8 | pass@32 |
|---|---|---|---|
| RFT (shortest-correct) | 0.571 [0.538, 0.605] | 0.910 [0.886,  ​​0.935] | 0.972 [0.955,​​0.990] |
| RFT + KL brake | 0.574 [0.541,​​0.607] | 0.908 [0.882,​​0.934] |​​0.967 [0.945,​​0.988] |
| GRPO |​​0.634 [0.598,​​0.669] |​​0.914 [0.887,​​0.940] |​​0.969 [0.951,​​0.988] |

GRPO's pass@1 CI does not overlap either RFT arm's,while at k=8 and k=32 all three overlap heavily. A paired bootstrap of the GRPO−RFT difference gives pass@1 gap of +0.058 [0.045,​​0.071] against RFT-shortest and +0.059 [0.046,​​0.073] against RFT+brake,both excluding zero,while pass@8/32 gaps(+0.005 to +0.020)all include zero.



**3. The gap survives restricting to correct solutions only**

Computing distinct-3 on verifier-correct GSM8K completions only:
- GRPO: 0.436 versus RFT arms: 0.512/0.509—a **15% gap** in the same direction as the all-completions result,on the same pass@32 that cannot tell the arms apart

This gap is **not a completion-length artifact**: RFT's correct completions are shorter(mean tokens: 272.7–272.8 versus 296.4). Truncating every completion to a shared word budget(157 words)gives 0.456 versus 0.531/0.528—a 14% gap,within one point of the untruncated 15%.



**4. Incorrect-answer diversity: a real but small effect**

$D_{\mathrm{incorrect}}$ (fraction of unique incorrect answers among incorrect completions only):0.758 (GRPO) versus 0.769/0.768 (RFT arms)—an order of magnitude smaller than the raw delta. Count-controlled variant(subsampling exactly 4 incorrect samples per problem)widens the gap to 0.873 versus 0.898/0.892,a 2.2–2.9% relative gap. Both versions are real but small;most of the raw magnitude in Table 1 is accuracy passing through the metric. The correct-only distinct-3 gap(14–15%)remains the strongest single diversity result.



**5. Hard benchmark rules out saturation**

On MATH-500 levels 4–5(pass@1 only 0.20–0.23,pass@16 below 0.67—substantial headroom):GRPO exceeds both RFT variants at pass@1 and pass@2,with a clustered bootstrap CI excluding zero at both, but from k=4 on the CI includes zero for both comparisons. This loss of separation happens far short of the benchmark's own ceiling,ruling out saturation as the explanation.



**6. No trained arm significantly improves over the starting checkpoint on hard problems**

Figure 2(a,b) shows the baseline sitting above every trained arm at every k. A clustered bootstrap of arm minus baseline finds:
- **RFT**: significantly below baseline at k=1,2,4(−3 to −3.5pp for both variants,CI excludes zero)—training actively cost coverage
- **GRPO**: never significantly differs from baseline,though point estimate drifts from −0.7pp at k=1 to −3.7pp at k=16

Read together with panel(c),GRPO's win over RFT on hard problems reflects a **smaller loss relative to the starting checkpoint** rather than a capability gain over it. This is a missing-control problem distinct from any failure of pass@k itself:pass@1 measured the right quantity,correctly finding more correct mass in GRPO than in RFT;the risk is entirely in reading a treatment gap as a capability claim with no control



**7. Self-distillation arm moves least**

At one seed each,length-penalized GRPO variants track plain GRPO closely on every diversity metric,while OPD is the outlier,nearly unchanged from baseline on every diversity metric and on pass@1(0.593 → 0.599,GSM8K). This is consistent with under-training or genuine diversity preservation;one seed cannot distinguish the two. OPD does not reproduce Kim et al.'s finding that self-distillation sharply lowers entropy,at this scale.



## Theoretical and Practical Implications

**Structural limitation of pass@k**: The paper establishes that pass@k is a statement about correct-answer mass;it says nothing about how post-training reorganized the model's output distribution to produce that mass. This is a construct-validity error:the metric is routinely used to certify diversity preservation,but structurally cannot detect it

**Two distinct claims**:
1. **Structural**: pass@k cannot see how probability mass is distributed across outputs,regardless of comparison
2. **Baseline comparison**: pass@1 correctly reports which treatment has more correct mass;the risk lies in reading that comparison without a control against the starting checkpoint

**Alignment tax interpretation**: The hard-problem regression is consistent with an alignment tax,where capability outside the training distribution degrades as a side effect of optimizing within it. The diversity narrowing measured elsewhere is a plausible,not causally established,mechanism for the same regression,consistent with Yue et al.'s boundary-narrowing result

**Practical guidance**: A practitioner reading only larger-k pass@k values would see no detectable separation past k=1–2;one reading entropy and answer-diversity would conclude the arms are opposites. These observations are not contradictory;they reflect different constructs. pass@k has its own blind spots,and treating large-k pass@k as sufficient evidence about diversity can lead to construct-validity errors when evaluating whether post-training changed more than the benchmark score shows

## Conclusion

The paper demonstrates that pass@k,at the k values the field reports,is not a valid standalone proxy for output diversity. Training Qwen2.5-1.5B-Instruct on grade-school math,GRPO and RFT move token entropy,answer entropy,and unique answers per prompt in opposite directions with zero seed overlap,yet pass@8 and pass@32 register no consistent winner. The correct-only lexical diversity among correct solutions is 15% lower for GRPO,holding after controlling for completion length;count-controlled checks isolate a small but real diversity gap among incorrect answers alone.



On a hard MATH-500 subset with real headroom,the arms separate only at low k and converge as k grows. Critically,no trained arm significantly improves over the starting checkpoint on hard problems:RFT is significantly worse,while GRPO is statistically indistinguishable from baseline—so GRPO's pass@1 edge over RFT there reflects a smaller loss,not a capability gain. This missing-control problem is distinct from any failure of pass@k itself.



**Future directions**: The paper suggests the need for evaluation protocols that explicitly measure diversity and capability retention,rather than relying on pass@k alone. Recent work like UCC@k [Lee et al,, 2026],which clusters correct completions by semantic similarity and sums per-cluster coverage,addresses part of this gap. The authors also highlight the need to scale these findings to 7B+ parameter models,and to isolate which mechanistic difference between training methods is responsible for the observed diversity changes.

---

_Markdown view of https://picx.dev/p/PlBn2A, served by PicX — AI-generated visual whiteboard summaries of research papers._
