Summary
Summary (Overview)
- This paper systematically studies the scaling properties of on-policy distillation (OPD) for transferring reinforcement learning (RL) expertise across model scales in math reasoning tasks.
- The authors identify a regular "useful-transfer" regime where gold score (G) rises approximately linearly with , the square root of token-level reverse KL divergence from the student initialization.
- Weak-to-strong transfer works remarkably well: compact RL experts can transfer capability to much larger students, with peak gold scores exceeding the teacher's own in every observed weak-to-strong pair.
- The paper fits joint power laws in student scale, teacher scale, and teacher gold score that predict peak gold score () and useful-transfer slope, extrapolating to held-out scales within one accuracy point.
- Two OPD variants are compared: Vanilla-OPD and Delta-OPD, with findings that bootstrapping weak-to-strong OPD does not improve on direct transfer, and off-policy cold starts harm weak-to-strong transfer.
Introduction and Theoretical Foundation
The paper addresses a fundamental question in LLM post-training: how much RL-acquired capability transfers across model scales, and how quickly?
Background: Modern reasoning models acquire task expertise through reinforcement learning. On-policy distillation (OPD) transfers this expertise between models—a teacher policy provides token-level supervision on rollouts sampled from the student, avoiding the exposure bias of offline teacher data.
Key Motivation: The central research question is whether OPD outcomes can be estimated from teacher and student scale before training, analogous to how reward model overoptimization studies characterize gold reward dynamics as functions of KL divergence from initialization.
Theoretical Foundation: The paper formalizes three distillation objectives:
- Vanilla-OPD: Minimizes sequence-level reverse KL divergence between policy and teacher:
- Delta-OPD: Derives token rewards from the policy shift the teacher acquired during RL:
- Off-policy distillation (OfPD): SFT on teacher demonstrations (used as a baseline/comparison).
Methodology
Models: Qwen2.5 Base models (not instruction-tuned) at 0.5B, 1.5B, 3B, 7B, and 14B parameters.
Training Setup:
- Teachers obtained via GRPO RL on mixed GSM8K + MATH training split (14.8K examples)
- Evaluation on corresponding mixed test split (6.3K examples)
- OPD and RL use the same training prompts
- Each OPD run: at most 10 epochs (580 updates) with periodic held-out evaluation
Experimental Design:
- 25 teacher–student combinations spanning weak-to-strong, same-base, and strong-to-weak setups
- Training progress metric: where is token-mean KL (square root of quadratic KL measure)
Key Comparisons:
- Pure OPD vs. Delta-OPD (17 matched pairs)
- Pure OPD vs. OfPD (off-policy supervision)
- Direct weak-to-strong OPD vs. bootstrapped transfer through intermediate scales
Empirical Validation / Results
1. Two-Phase Training Dynamics
OPD dynamics form a two-phase process:
- Useful-transfer regime: Gold score rises approximately linearly with at rate
- Noisy tail: Attenuated improvement, saturation, or regression (unlike the regular overoptimization regime of PPO)
2. Scaling Laws
Peak Gold Score ():
- Joint power laws in student parameters, effective teacher size, and measured teacher score
- Validation: conditioning on teacher score halves the leave-one-scale-out RMSE compared to scale-only baselines
- Extrapolates to held-out largest student/teacher within 0.7 (Vanilla-OPD) and 0.4 (Delta-OPD) accuracy points
Table 2: Validation of peak laws (accuracy points)
| Method | Scale-only RMSE | Joint law RMSE | Extrapolation error (largest held-out) |
|---|---|---|---|
| Vanilla-OPD | (baseline) | 0.5× | 0.7 |
| Delta-OPD | (baseline) | 0.5× | 0.4 |
3. Weak-to-Strong Transfer
- Every observed weak-to-strong pair: student's peak gold score exceeds teacher's own
- Margins shrink as teacher scale approaches student scale
- Example: 0.5B expert teaching 14B student shows substantial gains
4. Variant Comparisons
- Vanilla-OPD same-base peaks within 0.5 accuracy points of direct-RL reference at all five scales (exceeding at three)
- OfPD underperforms OPD in every cell: by 28.4 points in extreme weak-to-strong, <1 point in same-base cells
5. Off-Policy Cold Start Effects
- One epoch of SFT on teacher rollouts is increasingly harmful with weak-to-strong gap:
- 3B student: 6.6 points loss
- 7B student: 15.7 points loss
- 14B student: 19.4 points loss
- Root cause: cold start pins students near teacher's score, erasing 32 points of 14B student's initial capability
6. Bootstrapping Results
- Bootstrapped weak-to-strong OPD does not improve on direct transfer from the smallest expert
- Direct RL on the student remains above every weak-teacher variant
Theoretical and Practical Implications
Theoretical Contributions:
- Predictability: OPD outcomes can be estimated from student scale, teacher scale, and teacher score before training—enabling principled resource allocation
- Multiplicative error model: Scaling the student removes a constant fraction of whatever error teacher-induced supervision leaves, encoding scale-teacher interactions
- KL budget interpretation: The transfer extent () reads as an approximately scale-free KL budget rather than a scaling target
Practical Implications:
- Amortized expertise: RL expertise can be trained once at small scale and transferred predictably across a model family
- Design guidance:
- Pure on-policy supervision is critical (OfPD severely underperforms)
- Avoid off-policy SFT cold starts for weak-to-strong transfer
- Direct transfer beats bootstrapping through intermediate scales
- Cost savings: Enables choosing optimal teacher-student pairs without exhaustive experimentation
Conclusion
This work establishes that OPD scaling is predictable and regular in its initial phase, with gold score rising linearly in until a noisy tail. The fitted joint power laws in student scale, teacher scale, and teacher gold score enable accurate prediction of peak performance before training, with extrapolation errors under one accuracy point for the largest held-out configurations.
Key findings:
- Weak-to-strong transfer via OPD consistently exceeds teacher performance
- Delta-OPD and Vanilla-OPD show similar scaling behavior
- On-policy supervision is essential; off-policy cold starts are increasingly harmful with capability gaps
- Bootstrapping adds no value over direct transfer
Future Directions:
- Extending analysis beyond math reasoning to other domains
- Investigating whether the useful-transfer regime can be extended beyond its current duration
- Exploring adaptive KL budgets that maintain the linear regime longer
- Applying these scaling laws to other distillation objectives and architectures
Related papers
- Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Post-training updates leave behavioral shadows on unrelated inputs that can be extracted via single-word queries to transfer capabilities, yielding +5.34 points on HumanEval+.
- What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
Cross-harness credit assignment adds no detectable portability over within-harness grouping when exposure is held fixed, differing by only 0.25 pp on a held-out harness.
- Raven: The Harness of Harnesses for Composable Agentic Intelligence
Raven's Host Agent orchestrates model-harness pairs as composable units via DAG-based planning, beating Claude Code and Hermes Agent across coding, research, and design benchmarks.