Summary (Overview)

  • Core finding: pass@k, the standard evaluation metric for post-training quality, is structurally blind to how probability mass is distributed across distinct outputs. It measures the quantity of correct-answer mass, not its diversity or organization.
  • Empirical demonstration: Training Qwen2.5-1.5B-Instruct on GSM8K math problems with GRPO (Group Relative Policy Optimization) and RFT (Rejection-Sampling Fine-tuning) moves three complementary diversity measures (token entropy, answer entropy, unique answers per prompt)in opposite directions with zero overlap across three seeds per arm
  • Metric failure: pass@8 and pass@32 show no consistent winner between GRPO and RFT, despite large, seed-robust diversity differences. On a hard MATH-500 subset, arms separate only at low k and converge as k grows
  • Missing control problem: Compared against the starting checkpoint, no trained arm significantly improves hard-problem coverage. RFT is significantly worse than baseline, while GRPO is statistically indistinguishable from it, meaning GRPO's pass@1 edge over RFT reflects a smaller loss, not a capability gain
  • Key implication: An improved pass@k cannot certify that post-training preserved, rather than narrowed, what the starting checkpoint could do. The metric has a construct-validity gap that is structural, not an implementation artifact

Introduction and Theoretical Foundation

The paper addresses a fundamental question in post-training evaluation: how do we know whether reinforcement learning with verifiable rewards (RLVR) converts a broad-coverage base model into one that reliably solves tasks without secretly narrowing its capabilities?

The pass@k metric: The field's default protocol uses pass@k [Chen et al.,, 2021],the probability that at least one of k independently sampled attempts at a problem is correct, computed with the unbiased estimator:

1−(n−ck)(nk)1 - \frac{\binom{n-c}{k}}{\binom{n}{k}}

over n≥kn \geq k samples with c correct. pass@1 is treated as capability;pass@k at larger k is implicitly treated as a check on whether that capability cost the model's solution diversity.

Structural construct-validity gap: At the population level, for a problem with true per-sample correctness probability pp,:

pass@k(p)=1−(1−p)k\text{pass@k}(p) = 1 - (1-p)^k

a function of p alone. Two models with the same p have the same pass@k at every k, regardless of how the remaining probability mass is arranged across distinct correct solutions, incorrect ones, or reasoning paths. The paper states:

"pass@k is a statement about how much correct mass exists, not about how it is distributed. A post-training method that redistributes mass among already-correct outputs, or among incorrect ones, without changing p much, is therefore invisible to it by construction rather than by an implementation gap."

Related work and disagreements: The literature disagrees on whether RL narrows or broadens diversity:

  • Yue et al.[2025]: RLVR-trained models beat base models at small k but are overtaken once k reaches tens to hundreds, across six RL algorithms
  • Nicolicioiu et al.[2026]: On-policy self-distillation raises pass@1 while flattening pass@16,through a mechanism distinct from RL mode-seeking
  • Kim et al.[2026]: GRPO increases token entropy on two 7–8B models, while self-distillation decreases it—the reverse of what this paper reports at smaller scale

This paper reframes the contribution: rather than claiming "RL narrows diversity"(already disputed),it claims the protocol itself—pass@k at the k values the field reports—is not a valid standalone proxy for diversity,independent of which direction any method moves it.

Methodology

Model and data:

  • Base model: Qwen2.5-1.5B-Instruct (pre-post-training baseline, already instruction-tuned)
  • Training data: Frozen pool of 1,500 GSM8K problems [Cobbe et al.,, 2021], filtered to keep problems where 1–7 of 8 sampled rollouts land correct
  • Evaluation: In-domain on held-out 100 GSM8K problems(32 samples/problem); out-of-domain on MATH-500 [Lightman et al,, 2023] (16 samples/problem, never trained on)

Training arms:

  1. GRPO [Shao et al,, 2024]: group-relative advantage over 8 rollouts/prompt, clipped policy-gradient update, no KL penalty
  2. RFT (rejection-sampling fine-tuning): 8 rollouts/prompt, keep only the shortest verifier-passed one, fine-tuned with plain cross-entropy
  3. RFT + KL brake: adds Kullback-Leibler penalty toward the frozen starting policy
  4. Context arms (single seed each): GRPO with length penalties(λ = 0.2, 0.5)and on-policy self-distillation(OPD)

All generation uses temperature 1.0, top-p 1.0, no top-k truncation, 1,024-token budget. Each arm runs at three seeds(0, 1, 2),187 steps each

Metrics:

  • Token-level entropy, answer-level entropy, unique answers per prompt: computed on the policy model's own tokenizer output
  • pass@1/8/32: each seed's own value carries 95% bootstrap CI;3-seed summaries combine intervals with DerSimonian–Laird random-effects meta-analysis
  • Distinct-3 [Li et al,, 2016]: fraction of unique word-trigrams pooled across verifier-correct completions only, plus a length-truncated variant
  • DincorrectD_{\mathrm{incorrect}} : fraction of unique incorrect answers among incorrect completions only, plus a count-controlled variant(subsampling exactly 4 incorrect samples per problem)
  • Hard-problem evaluation: MATH-500 restricted to hardest levels(4–5,100 problems),pass@k for k ∈ {1,2,4,8,16},generated against each final checkpoint and the pre-intervention checkpoint

Empirical Validation / Results

1. Diversity moves in opposite, seed-robust directions

Table 1: GSM8K diversity metrics, change from first to last checkpoint, mean ± half-range across 3 seeds

ArmToken entropy ΔUnique answers/problem ΔAnswer entropy Δ
RFT (shortest-correct)+0.028 ± 0.003+0.56 ± 0.06+0.091 ± 0.017
RFT + KL brake+0.032 ± 0.001+0.52 ± 0.11+0.082 ± ​​0.009
GRPO-0.055 ± 0.004-1.30 ± 0.20-0.164 ± ​​0.018
Single seed each;context only:
GRPO + length penalty (λ=0.2)-0.055-1.26-0.163
GRPO + length penalty (λ=0.5)-0.050-1.20-0.150
On-policy self-distillation (OPD)-0.003-0.34-0.035

Both RFT arms increase all three measures in every seed;GRPO decreases all three in every seed,with no cross-arm overlap. Notably,GRPO narrows diversity despite being the arm least resembling self-distillation,contradicting a mode-seeking account.

2. pass@1 separates arms cleanly;pass@8/32 do not

Table 2: GSM8K pass@k, final checkpoint(3-seed DerSimonian–Laird pooled mean,95% CI)

Armpass@1pass@8pass@32
RFT (shortest-correct)0.571 [0.538, 0.605]0.910 [0.886, ​​0.935]0.972 [0.955,​​0.990]
RFT + KL brake0.574 [0.541,​​0.607]0.908 [0.882,​​0.934]​​0.967 [0.945,​​0.988]
GRPO​​0.634 [0.598,​​0.669]​​0.914 [0.887,​​0.940]​​0.969 [0.951,​​0.988]

GRPO's pass@1 CI does not overlap either RFT arm's,while at k=8 and k=32 all three overlap heavily. A paired bootstrap of the GRPO−RFT difference gives pass@1 gap of +0.058 [0.045,​​0.071] against RFT-shortest and +0.059 [0.046,​​0.073] against RFT+brake,both excluding zero,while pass@8/32 gaps(+0.005 to +0.020)all include zero.

3. The gap survives restricting to correct solutions only

Computing distinct-3 on verifier-correct GSM8K completions only:

  • GRPO: 0.436 versus RFT arms: 0.512/0.509—a 15% gap in the same direction as the all-completions result,on the same pass@32 that cannot tell the arms apart

This gap is not a completion-length artifact: RFT's correct completions are shorter(mean tokens: 272.7–272.8 versus 296.4). Truncating every completion to a shared word budget(157 words)gives 0.456 versus 0.531/0.528—a 14% gap,within one point of the untruncated 15%.

4. Incorrect-answer diversity: a real but small effect

DincorrectD_{\mathrm{incorrect}} (fraction of unique incorrect answers among incorrect completions only):0.758 (GRPO) versus 0.769/0.768 (RFT arms)—an order of magnitude smaller than the raw delta. Count-controlled variant(subsampling exactly 4 incorrect samples per problem)widens the gap to 0.873 versus 0.898/0.892,a 2.2–2.9% relative gap. Both versions are real but small;most of the raw magnitude in Table 1 is accuracy passing through the metric. The correct-only distinct-3 gap(14–15%)remains the strongest single diversity result.

5. Hard benchmark rules out saturation

On MATH-500 levels 4–5(pass@1 only 0.20–0.23,pass@16 below 0.67—substantial headroom):GRPO exceeds both RFT variants at pass@1 and pass@2,with a clustered bootstrap CI excluding zero at both, but from k=4 on the CI includes zero for both comparisons. This loss of separation happens far short of the benchmark's own ceiling,ruling out saturation as the explanation.

6. No trained arm significantly improves over the starting checkpoint on hard problems

Figure 2(a,b) shows the baseline sitting above every trained arm at every k. A clustered bootstrap of arm minus baseline finds:

  • RFT: significantly below baseline at k=1,2,4(−3 to −3.5pp for both variants,CI excludes zero)—training actively cost coverage
  • GRPO: never significantly differs from baseline,though point estimate drifts from −0.7pp at k=1 to −3.7pp at k=16

Read together with panel(c),GRPO's win over RFT on hard problems reflects a smaller loss relative to the starting checkpoint rather than a capability gain over it. This is a missing-control problem distinct from any failure of pass@k itself:pass@1 measured the right quantity,correctly finding more correct mass in GRPO than in RFT;the risk is entirely in reading a treatment gap as a capability claim with no control

7. Self-distillation arm moves least

At one seed each,length-penalized GRPO variants track plain GRPO closely on every diversity metric,while OPD is the outlier,nearly unchanged from baseline on every diversity metric and on pass@1(0.593 → 0.599,GSM8K). This is consistent with under-training or genuine diversity preservation;one seed cannot distinguish the two. OPD does not reproduce Kim et al.'s finding that self-distillation sharply lowers entropy,at this scale.

Theoretical and Practical Implications

Structural limitation of pass@k: The paper establishes that pass@k is a statement about correct-answer mass;it says nothing about how post-training reorganized the model's output distribution to produce that mass. This is a construct-validity error:the metric is routinely used to certify diversity preservation,but structurally cannot detect it

Two distinct claims:

  1. Structural: pass@k cannot see how probability mass is distributed across outputs,regardless of comparison
  2. Baseline comparison: pass@1 correctly reports which treatment has more correct mass;the risk lies in reading that comparison without a control against the starting checkpoint

Alignment tax interpretation: The hard-problem regression is consistent with an alignment tax,where capability outside the training distribution degrades as a side effect of optimizing within it. The diversity narrowing measured elsewhere is a plausible,not causally established,mechanism for the same regression,consistent with Yue et al.'s boundary-narrowing result

Practical guidance: A practitioner reading only larger-k pass@k values would see no detectable separation past k=1–2;one reading entropy and answer-diversity would conclude the arms are opposites. These observations are not contradictory;they reflect different constructs. pass@k has its own blind spots,and treating large-k pass@k as sufficient evidence about diversity can lead to construct-validity errors when evaluating whether post-training changed more than the benchmark score shows

Conclusion

The paper demonstrates that pass@k,at the k values the field reports,is not a valid standalone proxy for output diversity. Training Qwen2.5-1.5B-Instruct on grade-school math,GRPO and RFT move token entropy,answer entropy,and unique answers per prompt in opposite directions with zero seed overlap,yet pass@8 and pass@32 register no consistent winner. The correct-only lexical diversity among correct solutions is 15% lower for GRPO,holding after controlling for completion length;count-controlled checks isolate a small but real diversity gap among incorrect answers alone.

On a hard MATH-500 subset with real headroom,the arms separate only at low k and converge as k grows. Critically,no trained arm significantly improves over the starting checkpoint on hard problems:RFT is significantly worse,while GRPO is statistically indistinguishable from baseline—so GRPO's pass@1 edge over RFT there reflects a smaller loss,not a capability gain. This missing-control problem is distinct from any failure of pass@k itself.

Future directions: The paper suggests the need for evaluation protocols that explicitly measure diversity and capability retention,rather than relying on pass@k alone. Recent work like UCC@k [Lee et al,, 2026],which clusters correct completions by semantic similarity and sums per-cluster coverage,addresses part of this gap. The authors also highlight the need to scale these findings to 7B+ parameter models,and to isolate which mechanistic difference between training methods is responsible for the observed diversity changes.

Related papers