# On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

> Rollout policy in strong-to-weak distillation offers no consistent advantage; KL divergence direction and learning rate, not on-policy data, primarily determine performance and forgetting.

- **Source:** [arXiv](https://arxiv.org/abs/2609.35259)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/wp1aWy
- **Whiteboard:** https://picx.dev/p/wp1aWy/image

## Summary

## Summary (Overview)

- This paper presents a **controlled study of strong-to-weak knowledge distillation** that isolates the effect of rollout policy (on-policy vs. off-policy) from token-level KL divergence direction and learning rate, challenging the prevailing view that on-policy rollouts are inherently preferable.
- **Key finding**: On-policy rollouts offer **no consistent advantage** in final in-distribution accuracy, catastrophic forgetting, or parameter-update sparsity. Instead, **KL direction** most clearly shapes task performance and output coverage, while **learning rate** governs forgetting and update sparsity.
- The paper reveals an **objective-dependent asymmetry**: forward KL is remarkably robust to rollout policy changes (stable and strong performance across a student–teacher rollout-policy spectrum), whereas reverse KL is substantially more sensitive and favors student-generated rollouts.
- **On-policy data does help generalization** to harder task variants (Countdown-4E) under both KL directions, but this advantage **does not reliably persist after subsequent RLVR** (reinforcement learning with verifiable rewards).
- Findings are robust across model families (Llama 3, Qwen2.5), three reasoning domains (scientific, medical, arithmetic), and ablations (removing gradient clipping, sampled KL estimators, longer reasoning chains).

## Introduction and Theoretical Foundation

### Background and Motivation

Post-training of large language models (LLMs) uses several approaches: supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and knowledge distillation (KD). Strong-to-weak distillation—where a larger teacher transfers reasoning behavior to a smaller student—has become prominent in modern pipelines.

A key hypothesis in the literature is that the **policy used to generate training data** (rollout policy) is a central differentiating factor:
- **Off-policy methods** (e.g., SFT) learn from a fixed dataset or data-generating policy
- **On-policy methods** (e.g., RLVR) repeatedly train on outputs sampled from the model being optimized

Prior work has attributed several benefits to on-policy learning:
- Reduced/prevented **catastrophic forgetting** (Shenfeld et al., 2025; Chen et al., 2025)
- **Sparser parameter updates** (Mukherjee et al., 2025)
- Improved **generalization** (Chu et al., 2025; Zhang et al., 2026b; Yuan et al., 2026; Ming et al., 2025)

However, these comparisons typically vary multiple factors simultaneously (objective, reward signal, supervision density, optimization procedure), making it impossible to isolate the causal role of rollout policy.

### Theoretical Foundation: Distillation Objective

The distillation objective is formally defined as:

$$\mathcal{L}(\theta) = \mathbb{E}_{x \sim p_{data}} \mathbb{E}_{y \sim \rho(\cdot|x)} \left[ \frac{1}{L_y} \sum_{n=1}^{L_y} D\left( \pi^S_\theta(\cdot|x, y_{<n}) \| \pi^T(\cdot|x, y_{<n}) \right) \right] \tag{1}$$

where $\rho$ is the rollout policy, $\pi^S_\theta$ is the student, $\pi^T$ is the teacher, and $D$ is a divergence measure.

The two KL divergence directions are defined as:

**Forward KL** (mode-covering):
$$D_{F\text{-}KL}(\pi_S, \pi_T) = \mathbb{E}_{y_n \sim \pi_T}\left[\log \frac{\pi_T(y_n)}{\pi_S(y_n)}\right] = \sum_{v \in \mathcal{V}} \pi_T(v) \log \frac{\pi_T(v)}{\pi_S(v)} \tag{2}$$

**Reverse KL** (mode-seeking):
$$D_{R\text{-}KL}(\pi_S, \pi_T) = \mathbb{E}_{y_n \sim \pi_S}\left[\log \frac{\pi_S(y_n)}{\pi_T(y_n)}\right] = \sum_{v \in \mathcal{V}} \pi_S(v) \log \frac{\pi_S(v)}{\pi_T(v)} \tag{3}$$

The paper argues that rollout policy ($\rho$) and KL direction are **conceptually distinct** design choices. The conventional pairing (teacher rollouts + forward KL; student rollouts + reverse KL) arises from chain-rule decompositions of sequence-level KL, but this does not imply the crossed pairings are harder to optimize.

## Methodology

### Experimental Design

The paper uses strong-to-weak distillation as a **controlled testbed** where the data-generating policy can be changed while holding the remaining training pipeline fixed.

**Key controlled variables:**
- **Rollout policy**: OnPD (student-generated) vs. OffPD (teacher-generated)
- **Token-level KL direction**: Forward vs. Reverse
- **Learning rate**: $1 \times 10^{-5}$ or $5 \times 10^{-5}$ (plus a sweep from $1\times10^{-5}$ to $6\times10^{-5}$)

**Models and tasks:**
- Primary: Llama-3.1-8B teachers → Llama-3.2-1B students
- Secondary: Qwen2.5-7B teachers → Qwen2.5-1.5B students
- Tasks: **MedReason** (medical), **Science** (scientific), **Countdown-3** (arithmetic)
- Teachers are task-specifically fine-tuned (SFT + GRPO) and frozen during distillation

**Evaluation metrics:**
1. **Held-out ID accuracy** on the target task
2. **Catastrophic forgetting**: decrease in mean score across 7 OOD benchmarks (MMLU-Pro, TruthfulQA, HumanEval, IFEval, EQ-Bench, BBQ, ToxiGen)
3. **Parameter-update sparsity**: fraction of parameters with absolute change below $10^{-6}$

### Rollout-Policy Spectrum

To systematically vary the rollout policy, the authors define a likelihood-ratio-controlled policy $\pi_\lambda$:

$$(z_\lambda)_v = \frac{1}{2}(\log \pi_S(v) + \log \pi_T(v)) + \frac{\lambda}{2}(\log \pi_S(v) - \log \pi_T(v))$$

- $\lambda = -1$: recovers standard OffPD (teacher-favoured)
- $\lambda = 1$: recovers OnPD (student-favoured)
- $\lambda = 0$: symmetric midpoint
- $\lambda < -1$ or $\lambda > 1$: extrapolation beyond either policy

### Gradient Analysis

The paper derives token-level gradients for both KL directions:

**Forward KL gradient:**
$$\nabla_\theta D_{F\text{-}KL} = \sum_{v \in \mathcal{V}} (\pi^S_\theta(v) - \pi_T(v)) \nabla_\theta (z^S_\theta)_v \tag{4}$$

**Reverse KL gradient:**
$$\nabla_\theta D_{R\text{-}KL} = \sum_{v \in \mathcal{V}} \pi^S_\theta(v) \left[ \log \frac{\pi^S_\theta(v)}{\pi_T(v)} - D_{R\text{-}KL} \right] \nabla_\theta (z^S_\theta)_v \tag{5}$$

## Empirical Validation / Results

### Main Findings: No Consistent On-Policy Advantage

**Final task performance:**
- Best mean accuracies nearly identical: **72% for OnPD, 73% for OffPD**
- Forward KL achieves 71–73% across all rollout policies and learning rates
- Reverse KL ranges from 35% to 72%, substantially more sensitive to learning rate

**Catastrophic forgetting:**
- Mean OOD performance changes by at most **1.3 percentage points** at lower learning rate ($1\times10^{-5}$)
- Drops by **11.2–14.0 points** at higher learning rate ($5\times10^{-5}$)
- Learning rate, not rollout policy, is the dominant factor

**Parameter-update sparsity:**
- Sparsity ranges from **85.3–89.6%** at lower learning rate vs. **51.9–60.0%** at higher learning rate
- OffPD produces at least as sparse updates as OnPD in every matched comparison

### Rollout-Policy Spectrum Results

| Metric | Forward KL | Reverse KL |
|--------|-----------|------------|
| ID accuracy variation across λ | ≤5.2 pp (stays above 80%) | Large variation; benefits from λ>0 |
| Sensitivity to rollout policy | Robust | Highly sensitive |
| Learning at λ=1 (OnPD) | Effective despite high-entropy rollouts | Strong performance at LR $1\times10^{-5}$ |
| Stability at LR $5\times10^{-5}$ | Stable | Unstable, occasionally collapses |

### When On-Policy Data Helps

1. **Generalization to harder tasks**: On-policy rollouts ($\lambda > 0$) consistently yield **10–15% higher pass@k** accuracy on Countdown-4E (harder variant) under both KL directions.

2. **Output coverage**: Forward KL produces significantly higher pass@10 gains compared to reverse KL at comparable pass@1.

3. **Teacher-style transfer**: OnPD with reverse KL largely preserves the student's English response style (0% transfer to Spanish), whereas other configurations exhibit near-complete transfer (90.6–100%).

### RLVR After Distillation

- Reverse-KL checkpoints improve quickly but undergo **reward collapse**
- OffPD checkpoints with forward KL (LR $1\times10^{-5}$) improve steadily and reach **highest final performance**
- The initial generalization advantage of on-policy checkpoints **does not translate into a reliable advantage after RLVR**

### Robustness Checks

| Ablation | Result |
|----------|--------|
| Sampled KL (vs. full-vocabulary) | Same qualitative trends: forward KL robust, reverse KL sensitive, LR governs forgetting |
| No gradient clipping | Reverse KL deteriorates substantially; forward KL remains effective; LR still dominates forgetting/sparsity |
| Longer rollouts (Numina-MATH, ~622 tokens) | OnPD with reverse KL achieves highest MATH-500 accuracy (suggestive, single run); LR still governs forgetting/sparsity |

## Theoretical and Practical Implications

### Theoretical Contributions

The paper provides formal theoretical results explaining the empirical asymmetry:

**Theorem C.1 (Forward-KL rollout stability):** Under a bounded logit Jacobian condition, the forward-KL semi-gradient is uniformly Lipschitz in the rollout-induced distribution over prefixes:

$$\left\| \bar{\nabla}_\theta \mathcal{L}_{D_{F\text{-}KL}}(\theta; \rho) - \bar{\nabla}_\theta \mathcal{L}_{D_{F\text{-}KL}}(\theta; \rho') \right\|_2 \leq 2\sqrt{2}B \cdot \mathbb{E}_{x \sim p_{data}}[\text{TV}(\rho(\cdot|x), \rho'(\cdot|x))]$$

**Theorem C.2 (No distribution-free reverse-KL bound):** There is **no finite constant** $C$ independent of student/teacher distributions for which a similar bound holds for reverse KL. A small rollout change can produce arbitrarily large gradient changes when teacher and student assign very different probabilities to tokens.

**Corollary C.3:** A reverse-KL stability bound is recovered only under a bounded log-ratio range condition:

$$\max_{v \in \mathcal{V}} \log \frac{\pi^S_\theta(v|h_n)}{\pi_T(v|h_n)} - \min_{v \in \mathcal{V}} \log \frac{\pi^S_\theta(v|h_n)}{\pi_T(v|h_n)} \leq R$$

### Practical Implications

1. **Cost-benefit reconsideration**: On-policy distillation requires continual generation throughout training (computationally expensive), while off-policy methods can reuse existing data. Given the limited benefits observed, practitioners should weigh this additional cost carefully.

2. **Hyperparameter importance**: Learning rate is the primary determinant of catastrophic forgetting and update sparsity—more so than rollout policy. This echoes prior findings that learning rate strongly affects learning–forgetting trade-offs.

3. **Objective choice matters more**: The choice of token-level KL divergence (forward vs. reverse) has a larger effect on task performance, training stability, and output coverage than rollout policy.

4. **Recommendation**: The authors encourage future work on on-policy distillation to **include OffPD as a standard baseline**.

## Conclusion

### Main Takeaways

1. **Rollout policy does not play a central role** in strong-to-weak distillation for final in-distribution accuracy, catastrophic forgetting, or parameter-update sparsity.

2. **KL direction** more clearly shapes task performance and output coverage, with forward KL being more robust and providing better pass@k gains.

3. **Learning rate** governs catastrophic forgetting and update sparsity, with higher learning rates causing more forgetting and denser updates.

4. **On-policy rollouts do help generalization** to harder task variants, but this advantage does not reliably persist after subsequent RLVR.

5. **Forward KL is remarkably robust to rollout policy** (supported by both theory and experiments), while **reverse KL is substantially more sensitive** and favors student-generated rollouts.

### Limitations and Future Work

- The controlled setting focuses on student models ≤1.5B parameters and reasoning traces ≤2,000 tokens
- Extending to larger models and longer rollouts is a promising direction
- The teacher is kept fixed; future work could investigate student–teacher combinations
- The paper suggests that differences between SFT and RL cannot be attributed to rollout policy alone, motivating closer study of the learning objective, reward signal, supervision density, and optimization procedure

---

_Markdown view of https://picx.dev/p/wp1aWy, served by PicX — AI-generated visual whiteboard summaries of research papers._
