Summary (Overview)
- Problem: Multi-teacher on-policy distillation (MOPD) routes each prompt to a domain-specialist teacher, but fails to calibrate how strongly each teacher's feedback influences the shared student. This leads to imbalanced gradients where instruction-following feedback dominates and mathematics skills fail to transfer.
- Key Finding: In Qwen3.5 models at 9B, 4B, and 2B scales, label-routed MOPD does not outperform a student trained by the best single specialist, and retains only 14–32% of the mathematics expert's gain at an 8K evaluation budget.
- Proposed Solution: Domain-Normalized MOPD (DN-MOPD) rescales each domain's distillation advantages by the ratio of pooled to per-domain log-ratio standard deviation, using a bounded multiplier .
- Results: DN-MOPD improves the six-task average over MOPD by 1.17–2.36 percentage points at 16K and 2.47–3.08 at 8K evaluation budgets, across three random seeds and all model sizes.
- Attribution: Controls show the gain comes primarily from reducing the influence of dispersed instruction-following feedback rather than amplifying mathematics alone.
Introduction and Theoretical Foundation
Background and Motivation
Post-training of large language models increasingly relies on domain-specific reinforcement learning (RL) with verifiable rewards: mathematical reasoning (verifiable rewards), code execution feedback, and constraint checking for instruction following. These pipelines differ in data, rewards, and optimization, so they are typically run separately from a shared initialization, producing specialists that each excel in one domain.
Multi-teacher on-policy distillation (MOPD) integrates these specialists in policy space: the student samples responses, each prompt is routed to its domain expert, and that expert's token-level log-probabilities on the student's own response provide a dense distillation signal. MOPD has been adopted in the post-training of several frontier models (DeepSeek-AI, NVIDIA, Xiaomi, etc.).
The Core Problem: Feedback Imbalance
The routing rule specifies which expert supervises a prompt, but not how strongly its feedback moves the shared student. The paper's empirical investigation reveals:
- Token-level teacher–student log-ratios from the instruction-following expert are 2.3 to 4.4 times as dispersed as the pooled batch signal
- Mathematics feedback is about half as dispersed as the pooled signal
- For the initial 4B student, the instruction-following loss supplies 94% of the combined gradient under equal weights
"Routing determines where feedback comes from, but not how much it counts."
Theoretical Foundation: Distillation Advantages
The token-level distillation advantage is defined as:
where is the domain specialist teacher, is the student, and is the preceding context. A positive encourages the sampled token; a negative discourages it. The magnitude weights that token's policy-gradient contribution.
Methodology
Domain Normalization Rule
For each batch, DN-MOPD measures the spread of teacher–student log-ratios within each domain and rescales:
- Let be the token-level log-ratio (teacher minus student rollout log-probability)
- Compute = population standard deviation of over valid response tokens with domain
- Compute over all valid response tokens pooled across domains
- Apply the multiplier:
For a non-degenerate domain whose ratio is not clipped, the scaled rollout signal satisfies:
This amplifies feedback with smaller spread (e.g., mathematics) and attenuates feedback with larger spread (e.g., instruction following).
Student Update
The actor recomputes pre-update student log-probabilities on sampled responses, forming . The scaled advantage is detached and used in the clipped OPD loss:
where is the importance ratio, and , are policy-ratio clipping bounds.
Algorithm 1: DN-MOPD, one training iteration
Require: student π_u, teachers {T_d}, labeled prompts
1: Sample responses y_i ~ π_u(· | x_i)
2: Score each y_i with its domain teacher T_{d_i}
3: r_{i,t} ← ℓ_{i,t}^T - ℓ_{i,t}^roll on valid tokens
4: σ_all ← Std({r_{i,t}})
5: for each domain d do
6: σ_d ← Std({r_{i,t} : d_i = d})
7: w_d ← clip(σ_all / σ_d, 0.25, 4)
8: Use w_d = 1 if either statistic is degenerate
9: end for
10: Recompute A_{i,t} ← ℓ_{i,t}^T - ℓ_{i,t}^actor
11: Ã_{i,t} ← stopgrad(w_{d_i} A_{i,t})
12: Update π_u on à with the clipped OPD objective
Key Design Choices
- Sign preservation: Multiplies by a positive factor without subtracting the domain mean, preserving whether each advantage encourages or discourages the sampled token
- Bounded adjustment: The [0.25, 4] bounds limit extreme adjustments (the IF multiplier typically sits at the 0.25 floor)
- Zero additional cost: Reuses log-ratios already available from rollout scoring—no additional teacher call, teacher model, or learned router
- MOPD as special case: Setting every recovers standard label-routed MOPD
Empirical Validation / Results
Experimental Setup
- Models: Qwen3.5 at 9B, 4B, and 2B with independently trained specialist pools (mathematics, code, instruction following)
- Evaluation: Six public benchmarks—AIME25/AIME26 (math), LiveCodeBench v5/v6 (code), IFEval/IFBench (instruction following); sampling 64/6/16 responses per question respectively
- Training: 80 updates, student seed 42 (seeds 43, 44 for robustness), 8,192-token training response cap; evaluation at both 8K and 16K token caps
Main Results
Table 1: Capability integration at 9B (selected rows)
| Method | Math (avg@64) | Code (avg@6) | IF (avg@16) | Total |
|---|---|---|---|---|
| Initial student | 60.2 | 53.2 | 58.1 | 57.1 |
| Label-routed MOPD | 59.3 | 54.7 | 61.2 | 58.4 |
| DN-MOPD (Ours) | 63.3 | 53.9 | 61.6 | 59.6 |
Table 2: Capability integration at smaller scales (Total scores)
| Method | 4B Total | 2B Total |
|---|---|---|
| Initial student | 47.7 | 24.0 |
| Label-routed MOPD | 50.3 | 26.6 |
| DN-MOPD (Ours) | 52.5 | 29.0 |
Key Findings
- DN-MOPD consistently beats Label-routed MOPD at every size and both evaluation budgets, with paired confidence intervals above zero
- The main gain is in mathematics: At 16K, Label shows no math gain over the initial student at any size, while DN-MOPD improves math at every size
- DN-MOPD exceeds the strongest single-teacher student at every size (though some intervals include zero)
- Gains are not from longer generations: DN-MOPD produces shorter math answers with fewer responses hitting the generation cap (Table 4)
Ablation and Control Experiments
Table 3: Changes to label-routed OPD (Total point differences vs. Label at 16K)
| Variant | Teacher rule | Signal change | 9B | 4B | 2B |
|---|---|---|---|---|---|
| Uniform pool | Mixture | Unchanged | −0.41 | −0.69 | +0.30 |
| Dynamic router | Response | Unchanged | −0.25 | +0.28 | −0.07 |
| Annealed injection | Domain | Early imitation | −0.24 | +1.11 | +0.30 |
| Math ×2 only | Domain | Math scale | — | +1.20 | +0.27 |
| IF ×0.25 only | Domain | IF scale | — | +1.79 | +2.19 |
| DN-MOPD | Domain | Domain scale | +1.17 | +2.24 | +2.36 |
Key attribution: Reducing IF feedback alone recovers most of DN-MOPD's improvement (+1.79 at 4B, +2.19 at 2B), while amplifying math alone recovers only about half at 4B and little at 2B.
Training Duration Analysis
Table 5: Training-duration comparison (Total at 16K)
| Method | 80 updates | 160 updates | Δ Total |
|---|---|---|---|
| 9B: Label / DN-MOPD | 58.4 / 59.6 | 59.4 / 60.7 | +1.0 / +1.2 |
| 4B: Label / DN-MOPD | 50.3 / 52.5 | 51.6 / 53.1 | +1.3 / +0.5 |
| 2B: Label / DN-MOPD | 26.6 / 29.0 | 28.3 / 30.0 | +1.7 / +1.1 |
DN-MOPD remains ahead at every size when training continues to 160 updates, though the margin narrows at smaller scales.
Feedback Scale Diagnostics
- IF log-ratios are 2.3–4.4 times more dispersed than the pooled signal; math log-ratios about half as dispersed
- DN-MOPD amplifies mathematics by 1.5–2.1× at every size and keeps code near 1 at 4B and 2B
- The IF multiplier sits at the 0.25 floor in most batches, so clipping bounds rather than equalizes that domain's scale
- IF supplies about 1% of response tokens but up to half of the pooled variance
Theoretical and Practical Implications
Theoretical Implications
-
Teacher assignment is necessary but not sufficient for multi-teacher distillation. The scale of each teacher's feedback is an independent design choice that significantly affects integration quality.
-
Feedback dispersion, not just mean magnitude, shapes the shared gradient. Existing methods (e.g., Open-MOPD) allocate budget by mean absolute rewards or remaining capability gaps; DN-MOPD shows that dispersion-based calibration is a complementary and effective signal.
-
The gradient dominance of a single domain can be extreme: IF contributes 94% of the combined gradient for the initial 4B student under equal weights, despite representing only ~1% of tokens. This suggests that standard loss averaging can severely misallocate learning signal.
Practical Implications
-
Zero-cost improvement: DN-MOPD adds one operation to MOPD with no additional teacher calls, teacher models, or learned routers—making it a drop-in replacement for existing MOPD pipelines.
-
Fixed weights vs. adaptive estimation: Fixed weights near DN-MOPD's measured multipliers perform comparably at 9B and 4B, but per-batch estimation adds +1.10 points at 2B. This suggests practitioners can start with fixed weights and add adaptive estimation for smaller models.
-
Evaluation against single-teacher baselines: The paper shows MOPD students should be evaluated against the strongest single-teacher student, not just the initial model—a stronger reference that multi-teacher methods must beat.
Conclusion
Main Takeaways
- MOPD with label routing alone fails to transfer specialist skills effectively: it does not beat the strongest single-teacher student and loses most of the mathematics gain.
- DN-MOPD adds the missing control: rescaling each domain's distillation advantages by the batch-wise spread of teacher–student log-ratios, with bounded multipliers.
- The improvement is consistent and robust: across three model sizes, three student seeds, two evaluation budgets, and extended training durations.
- The mechanism is understood: DN-MOPD primarily works by limiting the influence of dispersed instruction-following feedback, which otherwise dominates the shared gradient.
Limitations and Future Directions
- Single model family: Comparisons use one expert pool per size within Qwen3.5; results may not generalize to other architectures or teacher configurations
- Fixed clipping bounds: The [0.25, 4] bounds were not tuned per size; the IF multiplier usually sits at the lower bound, so the rule bounds rather than equalizes that domain's scale
- Seed coverage: Three student seeds do not cover retraining the experts themselves
- Convergence behavior: The eventual ordering at convergence remains unresolved, though DN-MOPD remains ahead through 160 updates
Future Work
- Investigating whether adaptive clipping bounds or domain-specific normalization schemes yield further gains
- Applying DN-MOPD to other teacher pools, model families, and domain configurations
- Exploring the interaction between feedback scale calibration and other design choices (token allocation, reward refresh, etc.)
Related papers
- Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Post-training updates leave behavioral shadows on unrelated inputs that can be extracted via single-word queries to transfer capabilities, yielding +5.34 points on HumanEval+.
- Musec: MomentUm SpEctral Clipping for Stable Muon-type Training
Musec replaces Muon's spectral flattening with spectral clipping, achieving the first convergence guarantees for Muon-type optimizers in nonconvex nonsmooth settings with optimal complexity while stabilizing training across learning rates.
- Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems
AgentX-Model's dual-agent framework autonomously conducts long-horizon recommender model research, achieving 88% success across 636 experiments and 10-15% gains in production A/B tests.