# Gains and Collapse in On-Policy Distillation: A Reinforcement Learning Perspective

> On-policy distillation improves sampling efficiency without expanding capability, and its collapse stems from reward hacking when teacher preferences misalign with response quality.

- **Source:** [arXiv](https://arxiv.org/abs/2610.03185)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/Bg7bfh
- **Whiteboard:** https://picx.dev/p/Bg7bfh/image

## Summary

# On-Policy Distillation: Gains and Collapse from a Reinforcement Learning Perspective

## Summary (Overview)

- **Core contribution**: This paper provides a reinforcement learning (RL) interpretation of on-policy distillation (OPD), showing that the teacher acts as an *implicit reward model* that scores student-generated trajectories rather than supplying training trajectories.

- **Key finding on gains**: OPD improves performance **without expanding the student's capability boundary**—it makes correct responses that the initial student can already generate easier to sample, with gains diminishing as sampling budget increases.

- **Key finding on collapse**: OPD collapse is identified as **reward hacking**—the student faithfully fits the teacher's preference signal even when those preferences are misaligned with response quality, leading to overlong and repetitive generations.

- **Mitigation strategies**: Masking pathological (truncated) responses during training and using SFT initialization both effectively mitigate collapse, improving average accuracy by up to 3.08 percentage points with masking alone.

- **Practical implication**: The reliability of the teacher as an *evaluator* of student rollouts matters more than its generation quality—a teacher can generate high-quality text yet assign high rewards to degenerate student behaviors.

---

## Introduction and Theoretical Foundation

### Background and Motivation

On-policy distillation (OPD) has become a cornerstone of large-scale language model post-training pipelines, adopted by systems such as DeepSeek, Kimi, and others. In OPD, the current student samples its own responses, and the teacher supervises via token-level KL divergence at student-generated prefixes. While OPD can improve student performance, recent studies report unstable training and degenerate behaviors including excessive response length and repetition.

### Theoretical Foundation: OPD as Implicit Reward Optimization

The paper formulates the sampled-token reverse-KL variant of OPD as policy optimization with a teacher-derived reward.

**The conditional reverse-KL objective** is:

$$
\mathcal{L}_{T}(\theta ; \bar{S}) = \mathbb{E}_{h \sim d^{\bar{S}}} \left[ D_{\mathrm{KL}}\left(\pi_{S_{\theta}}(\cdot | h) \| \pi_{T}^{(\tau_{T})}(\cdot | h)\right) \right], \tag{1}
$$

where:
- $h = (x_q, y_{<t})$ is the prefix consisting of the prompt and previously generated tokens
- $d^{\bar{S}}$ is the prefix distribution induced by the rollout policy $\bar{S}$
- $\tau_T$ is the teacher scoring temperature
- $T$ is the fixed teacher, $S_\theta$ the trainable student

**Reward interpretation**: By expanding the KL divergence, the objective becomes a maximum-entropy policy objective:

$$
\mathcal{L}_{T}(\theta ; \bar{S}) = - \mathbb{E}_{h \sim d^{\bar{S}}} \left[ \mathbb{E}_{a \sim \pi_{S_{\theta}}(\cdot | h)} [ r_{T}(h, a) ] + \mathcal{H}(\pi_{S_{\theta}}(\cdot | h)) \right], \tag{2}
$$

where $r_T(h, a) = \log \pi_T^{(\tau_T)}(a \mid h)$ is the **token-level teacher reward** and $\mathcal{H}$ denotes entropy.

**Policy-gradient update**: The advantage for each rollout token is:

$$
A_{t}^{T} = r_{T}(h_{t}, y_{t}) - \log \pi_{\bar{S}}(y_{t} \mid h_{t}). \tag{3}
$$

The unclipped estimator satisfies:

$$
\mathbb{E}_{a \sim \pi_{\bar{S}}(\cdot | h)} \left[ A^{T}(h, a) \nabla_{\theta} \log \pi_{S_{\theta}}(a \mid h) \right] \Big|_{\theta = \bar{\theta}} = - \left. \nabla_{\theta} D_{\mathrm{KL}}\left(\pi_{S_{\theta}}(\cdot \mid h) \| \pi_{T}^{(\tau_{T})}(\cdot \mid h)\right) \right|_{\theta = \bar{\theta}}. \tag{4}
$$

**Key insight**: The teacher score $r_T$ and the update weight $A_t^T$ are distinct—the latter includes the student log-probability term from entropy regularization. When most advantages are negative, the main effect is to reduce token probabilities by different amounts, with more negative advantages creating stronger pressure to reduce probability.

---

## Methodology

### Experimental Settings

The paper studies three teacher–student settings:

| Teacher | Student initialization | Training data |
|---------|----------------------|---------------|
| JustRL-1.5B | DeepSeek-R1-Distill-Qwen-1.5B | DeepMath-103K (≥ 6) |
| DeepScaleR-1.5B-Preview | DeepSeek-R1-Distill-Qwen-1.5B | DeepMath-103K (≥ 6) |
| Qwen3-4B | Qwen3-1.7B-Base | OpenThoughts3 |

Key training details:
- **Rollout limits**: 16,384 tokens (JustRL/DeepScaleR settings), 8,192 tokens (Qwen3-4B setting)
- **Training**: 100 rollout iterations using the Slime framework, reference-KL coefficient of 0.001
- **Evaluation**: AIME24–26 benchmarks, mean@4 for routine evaluation, 256 responses per problem for coverage analyses

### Analytical Methods

1. **Pass@k curves**: Compare initial ($S_{\mathrm{INIT}}$) and trained ($S_{\mathrm{OPD}}$) students across sampling budgets (k ∈ {1, 2, 4, 8, 16, 32, 64, 128, 256} on AIME, up to k=1024 on AMC23)

2. **Coverage audit**: For problems solved by $S_{\mathrm{OPD}}$ but not $S_{\mathrm{INIT}}$ in 256-response pools, draw 768 additional responses from $S_{\mathrm{INIT}}$ and manually review derivations

3. **Min-10NN distance**: Measures response concentration (averages ten smallest nearest-neighbor distances among responses to the same problem; lower values indicate greater similarity)

4. **Selection gap (Γ)**: The mean ΔNLL of the least-preferred 10% minus that of the most-preferred 10% of responses (ranked by teacher advantage), measuring whether training favors teacher-preferred responses

5. **Advantage decile analysis**: Rank all sampled $S_{\mathrm{INIT}}$ responses by mean teacher advantage into ten groups (A1–A10) to characterize what the teacher rewards

---

## Empirical Validation / Results

### Finding 1: OPD Improves Sampling Efficiency Without Expanding Capability

**Pass@k results**: Under both JustRL-1.5B and DeepScaleR-1.5B-Preview teachers, $S_{\mathrm{OPD}}$ clearly outperforms $S_{\mathrm{INIT}}$ at small k (e.g., k=1), but $S_{\mathrm{INIT}}$ steadily catches up as k increases. At maximum k, pass@k equals observed coverage—the fraction of problems solved at least once.

**Coverage audit results**: After additional sampling (to 1024 responses) and manual review, **no problem remains solved only by $S_{\mathrm{OPD}}$**. The audited $S_{\mathrm{INIT}}$ set contains 73 solvable problems; $S_{\mathrm{OPD}}$ covers 65 (JustRL) and 63 (DeepScaleR).

**Per-problem success rates**: OPD increases per-problem success rates for **90.6%** of problems already solvable by $S_{\mathrm{INIT}}$ (JustRL teacher) and **84.4%** (DeepScaleR teacher), with mean@256 increases of **22.6%** and **14.7%** respectively.

### Finding 2: Collapse Is Reward Hacking, Not Optimization Failure

**Normal convergence**: The loss and advantage in the collapsed Qwen3-4B setting converge as smoothly as in the successful JustRL setting, with no qualitative difference in curve shapes.

**Response concentration**: Min-10NN distance keeps decreasing during training in both settings—mildly in JustRL, sharply in Qwen3-4B—indicating both students shrink toward narrower behavior spaces.

**Familiarity with $S_{\mathrm{INIT}}$**: In both settings, the trained student's responses lie in the low NLL region of $S_{\mathrm{INIT}}$, confirming that training exploits existing behaviors rather than discovering novel ones.

**Endpoint behaviors**:

| Model | Acc. (%) | Len. (k) | Trunc. (%) | Rep. (%) |
|-------|----------|----------|------------|----------|
| **JustRL-1.5B → DS-Distill-1.5B** | | | | |
| Teacher | 39.3 | 9.30 | 13.8 | 0.0 |
| $S_{\mathrm{INIT}}$ | 20.9 | 11.19 | 12.5 | 0.0 |
| $S_{\mathrm{OPD}}$ | 37.1 (+16.2) | 9.72 | 16.5 | 0.0 |
| **Qwen3-4B → Qwen3-1.7B-Base** | | | | |
| Teacher | 19.6 | 3.09 | 2.1 | 0.0 |
| $S_{\mathrm{INIT}}$ | 0.2 | 1.18 | 4.1 | 4.2 |
| $S_{\mathrm{OPD}}$ | 5.4 (+5.2) | 8.16 | **99.4** | **38.0** |

The Qwen3-4B setting exhibits severe collapse: 99.4% truncation rate and 38% repetition rate, far worse than the teacher (2.1% truncation, 0% repetition).

**Selection gap**: The JustRL setting's gap grows modestly from near zero to 0.07, while the Qwen3-4B setting's gap increases dramatically to **approximately 26**, showing overfitting to teacher-preferred pathological responses.

### Finding 3: The Teacher's Reward Signal Can Be Misaligned with Quality

**Advantage decile analysis**: In the JustRL setting, teacher advantage aligns with quality—correctness increases toward A10, length decreases toward the teacher's own average, no severe repetition. In the Qwen3-4B setting, the highest-advantage responses have near-zero correctness but substantially increased length, truncation, and repetition rates.

> **Critical insight**: The teacher itself rarely generates pathological responses (2.1% truncation, 0% repetition), yet assigns high advantage to such responses from the student. **Generation quality does not reliably reflect feedback quality on student rollouts.**

### Mitigation Results

**Masking intervention** (masking loss on responses reaching generation length limit):

| Setting | AMC23 | AIME24 | AIME25 | AIME26 | Avg. |
|---------|-------|--------|--------|--------|------|
| Base (pre-OPD) | 2.41 | 1.67 | 0.00 | 1.67 | 1.44 |
| 4B → Base | 29.82 | 3.33 | 3.33 | 4.17 | 10.16 |
| 4B → Base + mask | 34.64 | 8.33 | 5.83 | 4.17 | **13.24** |
| 8B → Base | 28.31 | 5.00 | 3.33 | 5.00 | 10.41 |
| 8B → Base + mask | 32.83 | 5.00 | 2.50 | 3.33 | 10.92 |
| 30B-A3B → Base | 28.61 | 6.67 | 2.50 | 4.17 | 10.49 |
| 30B-A3B → Base + mask | 33.13 | 10.00 | 4.17 | 5.00 | **13.07** |
| 4B → SFT (warmup) | 38.86 | 9.17 | 10.00 | 7.50 | **16.38** |

Masking improves average accuracy by **3.08, 0.51, and 2.58 percentage points** for the 4B, 8B, and 30B-A3B teachers respectively. The SFT warmup achieves the best final accuracy (16.38%), exceeding the Base-initialized OPD setting by 6.22 points.

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **OPD is fundamentally a selection mechanism, not a capability expansion mechanism**: The teacher's role is to reweight the student's existing response distribution according to its preferences. This reframes the theoretical understanding of what distillation achieves.

2. **Reward hacking can occur even with normal convergence**: The paper demonstrates that OPD collapse is not an optimization failure but rather a *successful optimization toward the wrong target*. This distinguishes it from explanations based on distribution mismatch or training instability.

3. **Teacher generation quality ≠ feedback quality**: A crucial theoretical insight is that a teacher's ability to generate high-quality text does not guarantee that its token-level scores on *student-generated* trajectories will align with quality. The reliability of the teacher as an *evaluator* must be examined independently.

4. **The advantage signal differs from the reward signal**: The paper clarifies that the teacher score $r_T$ and the update weight $A_t^T$ are distinct, with the latter incorporating student log-probability from entropy regularization—an important distinction for understanding OPD dynamics.

### Practical Implications

1. **Diagnostic tools**: The selection gap (Γ) and advantage decile analysis provide practical methods for detecting when OPD training is likely to collapse before it happens.

2. **Intervention strategies**: Masking pathological responses is a simple, effective remedy that requires no teacher changes. SFT warmup provides a complementary approach by improving the initial rollout distribution.

3. **Evaluation methodology**: The pass@k analysis across sampling budgets provides a robust way to distinguish capability expansion from sampling efficiency improvements—a distinction that matters for understanding what post-training actually achieves.

4. **Caveat on generality**: The paper notes that masking and SFT warmup are used primarily to test the explanation of collapse rather than proposed as general solutions, and findings are limited to small models and mathematical reasoning tasks.

---

## Conclusion

This paper provides a unified RL perspective on on-policy distillation, explaining both its successes and failures through the lens of the teacher as an implicit reward model. The key takeaways are:

1. **OPD amplifies what the student already samples**: It improves sampling efficiency on existing capabilities without expanding the solvable set.

2. **Collapse is reward hacking**: The student faithfully fits the teacher's preference signal even when that signal is misaligned with quality, converging normally while producing pathological outputs.

3. **The teacher's reliability as an evaluator is the critical factor**: The teacher's own generation quality does not guarantee reliable feedback on student rollouts.

4. **Mitigation is possible**: Masking unhealthy responses or using SFT initialization can prevent collapse while keeping the teacher fixed.

**Future directions** include extending these findings to larger models and other domains beyond mathematical reasoning, developing more general solutions for OPD training, and exploring whether larger sampling budgets might reveal additional correct responses from initial models.

---

_Markdown view of https://picx.dev/p/Bg7bfh, served by PicX — AI-generated visual whiteboard summaries of research papers._
