Full text not available for this paper

Summary (Overview)

  • This paper presents a controlled study of strong-to-weak knowledge distillation that isolates the effect of rollout policy (on-policy vs. off-policy) from token-level KL divergence direction and learning rate, challenging the prevailing view that on-policy rollouts are inherently preferable.
  • Key finding: On-policy rollouts offer no consistent advantage in final in-distribution accuracy, catastrophic forgetting, or parameter-update sparsity. Instead, KL direction most clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity.
  • The paper reveals an objective-dependent asymmetry: forward KL is remarkably robust to rollout policy changes (stable and strong performance across a student–teacher rollout-policy spectrum), whereas reverse KL is substantially more sensitive and favors student-generated rollouts.
  • On-policy data does help generalization to harder task variants (Countdown-4E) under both KL directions, but this advantage does not reliably persist after subsequent RLVR (reinforcement learning with verifiable rewards).
  • Findings are robust across model families (Llama 3, Qwen2.5), three reasoning domains (scientific, medical, arithmetic), and ablations (removing gradient clipping, sampled KL estimators, longer reasoning chains).

Introduction and Theoretical Foundation

Background and Motivation

Post-training of large language models (LLMs) uses several approaches: supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and knowledge distillation (KD). Strong-to-weak distillation—where a larger teacher transfers reasoning behavior to a smaller student—has become prominent in modern pipelines.

A key hypothesis in the literature is that the policy used to generate training data (rollout policy) is a central differentiating factor:

  • Off-policy methods (e.g., SFT) learn from a fixed dataset or data-generating policy
  • On-policy methods (e.g., RLVR) repeatedly train on outputs sampled from the model being optimized

Prior work has attributed several benefits to on-policy learning:

  • Reduced/prevented catastrophic forgetting (Shenfeld et al., 2025; Chen et al., 2025)
  • Sparser parameter updates (Mukherjee et al., 2025)
  • Improved generalization (Chu et al., 2025; Zhang et al., 2026b; Yuan et al., 2026; Ming et al., 2025)

However, these comparisons typically vary multiple factors simultaneously (objective, reward signal, supervision density, optimization procedure), making it impossible to isolate the causal role of rollout policy.

Theoretical Foundation: Distillation Objective

The distillation objective is formally defined as:

L(θ)=Ex∼pdataEy∼ρ(⋅∣x)[1Ly∑n=1LyD(πθS(⋅∣x,y<n)∥πT(⋅∣x,y<n))](1)\mathcal{L}(\theta) = \mathbb{E}_{x \sim p_{data}} \mathbb{E}_{y \sim \rho(\cdot|x)} \left[ \frac{1}{L_y} \sum_{n=1}^{L_y} D\left( \pi^S_\theta(\cdot|x, y_{<n}) \| \pi^T(\cdot|x, y_{<n}) \right) \right] \tag{1}

where ρ\rho is the rollout policy, πθS\pi^S_\theta is the student, πT\pi^T is the teacher, and DD is a divergence measure.

The two KL divergence directions are defined as:

Forward KL (mode-covering):

DF-KL(πS,πT)=Eyn∼πT[log⁡πT(yn)πS(yn)]=∑v∈VπT(v)log⁡πT(v)πS(v)(2)D_{F\text{-}KL}(\pi_S, \pi_T) = \mathbb{E}_{y_n \sim \pi_T}\left[\log \frac{\pi_T(y_n)}{\pi_S(y_n)}\right] = \sum_{v \in \mathcal{V}} \pi_T(v) \log \frac{\pi_T(v)}{\pi_S(v)} \tag{2}

Reverse KL (mode-seeking):

DR-KL(πS,πT)=Eyn∼πS[log⁡πS(yn)πT(yn)]=∑v∈VπS(v)log⁡πS(v)πT(v)(3)D_{R\text{-}KL}(\pi_S, \pi_T) = \mathbb{E}_{y_n \sim \pi_S}\left[\log \frac{\pi_S(y_n)}{\pi_T(y_n)}\right] = \sum_{v \in \mathcal{V}} \pi_S(v) \log \frac{\pi_S(v)}{\pi_T(v)} \tag{3}

The paper argues that rollout policy (ρ\rho) and KL direction are conceptually distinct design choices. The conventional pairing (teacher rollouts + forward KL; student rollouts + reverse KL) arises from chain-rule decompositions of sequence-level KL, but this does not imply the crossed pairings are harder to optimize.

Methodology

Experimental Design

The paper uses strong-to-weak distillation as a controlled testbed where the data-generating policy can be changed while holding the remaining training pipeline fixed.

Key controlled variables:

  • Rollout policy: OnPD (student-generated) vs. OffPD (teacher-generated)
  • Token-level KL direction: Forward vs. Reverse
  • Learning rate: 1×10−51 \times 10^{-5} or 5×10−55 \times 10^{-5} (plus a sweep from 1×10−51\times10^{-5} to 6×10−56\times10^{-5})

Models and tasks:

  • Primary: Llama-3.1-8B teachers → Llama-3.2-1B students
  • Secondary: Qwen2.5-7B teachers → Qwen2.5-1.5B students
  • Tasks: MedReason (medical), Science (scientific), Countdown-3 (arithmetic)
  • Teachers are task-specifically fine-tuned (SFT + GRPO) and frozen during distillation

Evaluation metrics:

  1. Held-out ID accuracy on the target task
  2. Catastrophic forgetting: decrease in mean score across 7 OOD benchmarks (MMLU-Pro, TruthfulQA, HumanEval, IFEval, EQ-Bench, BBQ, ToxiGen)
  3. Parameter-update sparsity: fraction of parameters with absolute change below 10−610^{-6}

Rollout-Policy Spectrum

To systematically vary the rollout policy, the authors define a likelihood-ratio-controlled policy πλ\pi_\lambda:

(zλ)v=12(log⁡πS(v)+log⁡πT(v))+λ2(log⁡πS(v)−log⁡πT(v))(z_\lambda)_v = \frac{1}{2}(\log \pi_S(v) + \log \pi_T(v)) + \frac{\lambda}{2}(\log \pi_S(v) - \log \pi_T(v))
  • λ=−1\lambda = -1: recovers standard OffPD (teacher-favoured)
  • λ=1\lambda = 1: recovers OnPD (student-favoured)
  • λ=0\lambda = 0: symmetric midpoint
  • λ<−1\lambda < -1 or λ>1\lambda > 1: extrapolation beyond either policy

Gradient Analysis

The paper derives token-level gradients for both KL directions:

Forward KL gradient:

∇θDF-KL=∑v∈V(πθS(v)−πT(v))∇θ(zθS)v(4)\nabla_\theta D_{F\text{-}KL} = \sum_{v \in \mathcal{V}} (\pi^S_\theta(v) - \pi_T(v)) \nabla_\theta (z^S_\theta)_v \tag{4}

Reverse KL gradient:

∇θDR-KL=∑v∈VπθS(v)[log⁡πθS(v)πT(v)−DR-KL]∇θ(zθS)v(5)\nabla_\theta D_{R\text{-}KL} = \sum_{v \in \mathcal{V}} \pi^S_\theta(v) \left[ \log \frac{\pi^S_\theta(v)}{\pi_T(v)} - D_{R\text{-}KL} \right] \nabla_\theta (z^S_\theta)_v \tag{5}

Empirical Validation / Results

Main Findings: No Consistent On-Policy Advantage

Final task performance:

  • Best mean accuracies nearly identical: 72% for OnPD, 73% for OffPD
  • Forward KL achieves 71–73% across all rollout policies and learning rates
  • Reverse KL ranges from 35% to 72%, substantially more sensitive to learning rate

Catastrophic forgetting:

  • Mean OOD performance changes by at most 1.3 percentage points at lower learning rate (1×10−51\times10^{-5})
  • Drops by 11.2–14.0 points at higher learning rate (5×10−55\times10^{-5})
  • Learning rate, not rollout policy, is the dominant factor

Parameter-update sparsity:

  • Sparsity ranges from 85.3–89.6% at lower learning rate vs. 51.9–60.0% at higher learning rate
  • OffPD produces at least as sparse updates as OnPD in every matched comparison

Rollout-Policy Spectrum Results

MetricForward KLReverse KL
ID accuracy variation across λ≤5.2 pp (stays above 80%)Large variation; benefits from λ>0
Sensitivity to rollout policyRobustHighly sensitive
Learning at λ=1 (OnPD)Effective despite high-entropy rolloutsStrong performance at LR 1×10−51\times10^{-5}
Stability at LR 5×10−55\times10^{-5}StableUnstable, occasionally collapses

When On-Policy Data Helps

  1. Generalization to harder tasks: On-policy rollouts (λ>0\lambda > 0) consistently yield 10–15% higher pass@k accuracy on Countdown-4E (harder variant) under both KL directions.

  2. Output coverage: Forward KL produces significantly higher pass@10 gains compared to reverse KL at comparable pass@1.

  3. Teacher-style transfer: OnPD with reverse KL largely preserves the student's English response style (0% transfer to Spanish), whereas other configurations exhibit near-complete transfer (90.6–100%).

RLVR After Distillation

  • Reverse-KL checkpoints improve quickly but undergo reward collapse
  • OffPD checkpoints with forward KL (LR 1×10−51\times10^{-5}) improve steadily and reach highest final performance
  • The initial generalization advantage of on-policy checkpoints does not translate into a reliable advantage after RLVR

Robustness Checks

AblationResult
Sampled KL (vs. full-vocabulary)Same qualitative trends: forward KL robust, reverse KL sensitive, LR governs forgetting
No gradient clippingReverse KL deteriorates substantially; forward KL remains effective; LR still dominates forgetting/sparsity
Longer rollouts (Numina-MATH, ~622 tokens)OnPD with reverse KL achieves highest MATH-500 accuracy (suggestive, single run); LR still governs forgetting/sparsity

Theoretical and Practical Implications

Theoretical Contributions

The paper provides formal theoretical results explaining the empirical asymmetry:

Theorem C.1 (Forward-KL rollout stability): Under a bounded logit Jacobian condition, the forward-KL semi-gradient is uniformly Lipschitz in the rollout-induced distribution over prefixes:

∥∇ˉθLDF-KL(θ;ρ)−∇ˉθLDF-KL(θ;ρ′)∥2≤22B⋅Ex∼pdata[TV(ρ(⋅∣x),ρ′(⋅∣x))]\left\| \bar{\nabla}_\theta \mathcal{L}_{D_{F\text{-}KL}}(\theta; \rho) - \bar{\nabla}_\theta \mathcal{L}_{D_{F\text{-}KL}}(\theta; \rho') \right\|_2 \leq 2\sqrt{2}B \cdot \mathbb{E}_{x \sim p_{data}}[\text{TV}(\rho(\cdot|x), \rho'(\cdot|x))]

Theorem C.2 (No distribution-free reverse-KL bound): There is no finite constant CC independent of student/teacher distributions for which a similar bound holds for reverse KL. A small rollout change can produce arbitrarily large gradient changes when teacher and student assign very different probabilities to tokens.

Corollary C.3: A reverse-KL stability bound is recovered only under a bounded log-ratio range condition:

max⁡v∈Vlog⁡πθS(v∣hn)πT(v∣hn)−min⁡v∈Vlog⁡πθS(v∣hn)πT(v∣hn)≤R\max_{v \in \mathcal{V}} \log \frac{\pi^S_\theta(v|h_n)}{\pi_T(v|h_n)} - \min_{v \in \mathcal{V}} \log \frac{\pi^S_\theta(v|h_n)}{\pi_T(v|h_n)} \leq R

Practical Implications

  1. Cost-benefit reconsideration: On-policy distillation requires continual generation throughout training (computationally expensive), while off-policy methods can reuse existing data. Given the limited benefits observed, practitioners should weigh this additional cost carefully.

  2. Hyperparameter importance: Learning rate is the primary determinant of catastrophic forgetting and update sparsity—more so than rollout policy. This echoes prior findings that learning rate strongly affects learning–forgetting trade-offs.

  3. Objective choice matters more: The choice of token-level KL divergence (forward vs. reverse) has a larger effect on task performance, training stability, and output coverage than rollout policy.

  4. Recommendation: The authors encourage future work on on-policy distillation to include OffPD as a standard baseline.

Conclusion

Main Takeaways

  1. Rollout policy does not play a central role in strong-to-weak distillation for final in-distribution accuracy, catastrophic forgetting, or parameter-update sparsity.

  2. KL direction more clearly shapes task performance and output coverage, with forward KL being more robust and providing better pass@k gains.

  3. Learning rate governs catastrophic forgetting and update sparsity, with higher learning rates causing more forgetting and denser updates.

  4. On-policy rollouts do help generalization to harder task variants, but this advantage does not reliably persist after subsequent RLVR.

  5. Forward KL is remarkably robust to rollout policy (supported by both theory and experiments), while reverse KL is substantially more sensitive and favors student-generated rollouts.

Limitations and Future Work

  • The controlled setting focuses on student models ≤1.5B parameters and reasoning traces ≤2,000 tokens
  • Extending to larger models and longer rollouts is a promising direction
  • The teacher is kept fixed; future work could investigate student–teacher combinations
  • The paper suggests that differences between SFT and RL cannot be attributed to rollout policy alone, motivating closer study of the learning objective, reward signal, supervision density, and optimization procedure

Related papers