Summary

Summary (Overview)

  • This paper systematically studies the scaling properties of on-policy distillation (OPD) for transferring reinforcement learning (RL) expertise across model scales in math reasoning tasks.
  • The authors identify a regular "useful-transfer" regime where gold score (G) rises approximately linearly with d=KL(πθ∥πref)d = \sqrt{KL(\pi_\theta \| \pi_{ref})}, the square root of token-level reverse KL divergence from the student initialization.
  • Weak-to-strong transfer works remarkably well: compact RL experts can transfer capability to much larger students, with peak gold scores exceeding the teacher's own in every observed weak-to-strong pair.
  • The paper fits joint power laws in student scale, teacher scale, and teacher gold score that predict peak gold score (GpeakG_{peak}) and useful-transfer slope, extrapolating to held-out scales within one accuracy point.
  • Two OPD variants are compared: Vanilla-OPD and Delta-OPD, with findings that bootstrapping weak-to-strong OPD does not improve on direct transfer, and off-policy cold starts harm weak-to-strong transfer.

Introduction and Theoretical Foundation

The paper addresses a fundamental question in LLM post-training: how much RL-acquired capability transfers across model scales, and how quickly?

Background: Modern reasoning models acquire task expertise through reinforcement learning. On-policy distillation (OPD) transfers this expertise between models—a teacher policy provides token-level supervision on rollouts sampled from the student, avoiding the exposure bias of offline teacher data.

Key Motivation: The central research question is whether OPD outcomes can be estimated from teacher and student scale before training, analogous to how reward model overoptimization studies characterize gold reward dynamics as functions of KL divergence from initialization.

Theoretical Foundation: The paper formalizes three distillation objectives:

  1. Vanilla-OPD: Minimizes sequence-level reverse KL divergence between policy and teacher:
min⁡θLV(θ)=min⁡θEx∼D,y∼πθ(⋅∣x)[KL(πθ(y∣x)∥πT(y∣x))](1)\min_{\theta} \mathcal{L}_{V}(\theta) = \min_{\theta} \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_{\theta}(\cdot|x)} \left[ \text{KL}(\pi_{\theta}(y|x) \| \pi_{T}(y|x)) \right] \tag{1}
  1. Delta-OPD: Derives token rewards from the policy shift the teacher acquired during RL:
AtΔ(x,y)=rtΔ:=log⁡πT(yt∣x,y<t)−log⁡πTbase(yt∣x,y<t)A_{t}^{\Delta}(x, y) = r_{t}^{\Delta} := \log \pi_{T}(y_t | x, y_{<t}) - \log \pi_{T}^{\text{base}}(y_t | x, y_{<t})
  1. Off-policy distillation (OfPD): SFT on teacher demonstrations (used as a baseline/comparison).

Methodology

Models: Qwen2.5 Base models (not instruction-tuned) at 0.5B, 1.5B, 3B, 7B, and 14B parameters.

Training Setup:

  • Teachers obtained via GRPO RL on mixed GSM8K + MATH training split (14.8K examples)
  • Evaluation on corresponding mixed test split (6.3K examples)
  • OPD and RL use the same training prompts
  • Each OPD run: at most 10 epochs (580 updates) with periodic held-out evaluation

Experimental Design:

  • 25 teacher–student combinations spanning weak-to-strong, same-base, and strong-to-weak setups
  • Training progress metric: d:=k3d := \sqrt{k_3} where k3k_3 is token-mean KL (square root of quadratic KL measure)

Key Comparisons:

  1. Pure OPD vs. Delta-OPD (17 matched pairs)
  2. Pure OPD vs. OfPD (off-policy supervision)
  3. Direct weak-to-strong OPD vs. bootstrapped transfer through intermediate scales

Empirical Validation / Results

1. Two-Phase Training Dynamics

OPD dynamics form a two-phase process:

  • Useful-transfer regime: Gold score rises approximately linearly with dd at rate mm
  • Noisy tail: Attenuated improvement, saturation, or regression (unlike the regular overoptimization regime of PPO)

2. Scaling Laws

Peak Gold Score (GpeakG_{peak}):

  • Joint power laws in student parameters, effective teacher size, and measured teacher score
  • Validation: conditioning on teacher score halves the leave-one-scale-out RMSE compared to scale-only baselines
  • Extrapolates to held-out largest student/teacher within 0.7 (Vanilla-OPD) and 0.4 (Delta-OPD) accuracy points

Table 2: Validation of peak laws (accuracy points)

MethodScale-only RMSEJoint law RMSEExtrapolation error (largest held-out)
Vanilla-OPD(baseline)0.5×0.7
Delta-OPD(baseline)0.5×0.4

3. Weak-to-Strong Transfer

  • Every observed weak-to-strong pair: student's peak gold score exceeds teacher's own
  • Margins shrink as teacher scale approaches student scale
  • Example: 0.5B expert teaching 14B student shows substantial gains

4. Variant Comparisons

  • Vanilla-OPD same-base peaks within 0.5 accuracy points of direct-RL reference at all five scales (exceeding at three)
  • OfPD underperforms OPD in every cell: by 28.4 points in extreme weak-to-strong, <1 point in same-base cells

5. Off-Policy Cold Start Effects

  • One epoch of SFT on teacher rollouts is increasingly harmful with weak-to-strong gap:
    • 3B student: 6.6 points loss
    • 7B student: 15.7 points loss
    • 14B student: 19.4 points loss
  • Root cause: cold start pins students near teacher's score, erasing 32 points of 14B student's initial capability

6. Bootstrapping Results

  • Bootstrapped weak-to-strong OPD does not improve on direct transfer from the smallest expert
  • Direct RL on the student remains above every weak-teacher variant

Theoretical and Practical Implications

Theoretical Contributions:

  1. Predictability: OPD outcomes can be estimated from student scale, teacher scale, and teacher score before training—enabling principled resource allocation
  2. Multiplicative error model: Scaling the student removes a constant fraction of whatever error teacher-induced supervision leaves, encoding scale-teacher interactions
  3. KL budget interpretation: The transfer extent (dtransferd_{transfer}) reads as an approximately scale-free KL budget rather than a scaling target

Practical Implications:

  1. Amortized expertise: RL expertise can be trained once at small scale and transferred predictably across a model family
  2. Design guidance:
    • Pure on-policy supervision is critical (OfPD severely underperforms)
    • Avoid off-policy SFT cold starts for weak-to-strong transfer
    • Direct transfer beats bootstrapping through intermediate scales
  3. Cost savings: Enables choosing optimal teacher-student pairs without exhaustive experimentation

Conclusion

This work establishes that OPD scaling is predictable and regular in its initial phase, with gold score rising linearly in d=KLd = \sqrt{KL} until a noisy tail. The fitted joint power laws in student scale, teacher scale, and teacher gold score enable accurate prediction of peak performance before training, with extrapolation errors under one accuracy point for the largest held-out configurations.

Key findings:

  • Weak-to-strong transfer via OPD consistently exceeds teacher performance
  • Delta-OPD and Vanilla-OPD show similar scaling behavior
  • On-policy supervision is essential; off-policy cold starts are increasingly harmful with capability gaps
  • Bootstrapping adds no value over direct transfer

Future Directions:

  • Extending analysis beyond math reasoning to other domains
  • Investigating whether the useful-transfer regime can be extended beyond its current duration
  • Exploring adaptive KL budgets that maintain the linear regime longer
  • Applying these scaling laws to other distillation objectives and architectures

Related papers