# Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability

> Expanding alignment coverage via span supervision hurts cross-tokenizer distillation; strict 1:1 token alignment on a compact top-16 shared vocabulary subset outperforms all complex baselines.

- **Source:** [arXiv](https://arxiv.org/abs/2610.08448)
- **Published:** 2026-10-08
- **Permalink:** https://picx.dev/p/tlUP8j
- **Whiteboard:** https://picx.dev/p/tlUP8j/image

## Summary

## Summary (Overview)

- **Core Question**: The paper investigates whether expanding alignment coverage in Cross-Tokenizer On-Policy Distillation (OPD) improves student learning, or whether supervision reliability matters more than coverage.
- **Key Finding 1**: Strict 1:1 token alignment already covers 85.57–96.98% of student-generated tokens across three heterogeneous teacher–student pairs, despite static vocabulary Jaccard overlap of only 39.49–64.87%.
- **Key Finding 2**: Adding span MSE supervision on mismatch groups achieves complete structural coverage but *reduces* accuracy across all 18 tested positive-weight settings (0.27–1.20 percentage points below strict baselines).
- **Key Finding 3**: A compact student-selected top-16 subset of the shared vocabulary retains at least 96% of the full-average improvement of full shared-vocabulary OPD, outperforming all evaluated cross-tokenizer baselines (ULD, Extended ULD, GOLD, SimCT).
- **Key Finding 4**: Gradient diagnostics reveal that span gradients have weak or negative directional agreement with strict gradients and grow in relative magnitude during training, explaining the accuracy drop from span supervision.

---

## Introduction and Theoretical Foundation

### Background

On-Policy Distillation (OPD) trains a student language model on its own generations using teacher feedback. The student samples trajectories from its current policy, and the teacher provides supervision at the exact prefixes the student visits. This connects teacher feedback to the student's evolving behavior.

### The Cross-Tokenizer Challenge

Extending OPD across model families introduces a fundamental alignment problem: different tokenizers assign different token boundaries and use different vocabularies. The same response can be tokenized differently, and next-token distributions are defined over different vocabulary spaces. Cross-Tokenizer OPD must therefore handle alignment at two levels:

1. **Sequence level**: Matching token boundaries between teacher and student tokenizations
2. **Vocabulary level**: Comparing predictions across different vocabulary spaces

### Key Concepts

**Alignment coverage** describes how much of the response and vocabularies can be compared across tokenizers. **Supervision reliability** concerns whether the resulting teacher targets provide useful guidance for student learning.

The paper challenges the assumption that maximizing alignment coverage necessarily improves distillation quality. Prior work (SimCT, Byte-Prefix Marginalization) has focused on recovering excluded supervision signals, but the authors question whether these recovered signals actually contribute to learning.

### Theoretical Foundation

With a shared tokenizer, OPD minimizes:

$$
\mathcal{L}_{\mathrm{OPD}}(\theta) = \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_{\theta}(\cdot | x)} \left[ \sum_{i = 1}^{L} \mathrm{KL}(\pi_{\theta}(\cdot | x, y_{< i}) \| \pi_{\mathrm{T}}(\cdot | x, y_{< i})) \right].\tag{1}
$$

This can be viewed as dense KL-constrained reinforcement learning, where the teacher distribution induces a token-level reward.

---

## Methodology

### Token-Group Alignment

For scoring and alignment, the decoded student response is tokenized with both models' tokenizers. The response is partitioned into R pairs of aligned token groups:

$$
\mathcal{S}(y) = \left\{\left(S_{r}^{\theta}, S_{r}^{\mathrm{T}}\right)\right\}_{r = 1}^{R}.\tag{2}
$$

Groups are divided into strict 1:1 groups and mismatch groups:

$$
\mathcal{A}_{1:1}(y) = \left\{r: |S_{r}^{\theta}| = |S_{r}^{\mathrm{T}}| = 1\right\}, \quad \mathcal{A}_{\text{mis}}(y) = \{1, \dots, R\} \setminus \mathcal{A}_{1:1}(y).\tag{3}
$$

### Shared Vocabulary Restriction

Both distributions are restricted and renormalized over the shared vocabulary $\mathcal{V}_{\cap}$:

$$
\bar{\pi}_{\theta}(w \mid x, y_{< i_{r}}) = \frac{\pi_{\theta}(w \mid x, y_{< i_{r}})}{\sum_{w^{\prime} \in \mathcal{V}_{\cap}} \pi_{\theta}(w^{\prime} \mid x, y_{< i_{r}})}, \quad \bar{\pi}_{\mathrm{T}}(w \mid x, v_{< j_{r}}) = \frac{\pi_{\mathrm{T}}(w \mid x, v_{< j_{r}})}{\sum_{w^{\prime} \in \mathcal{V}_{\cap}} \pi_{\mathrm{T}}(w^{\prime} \mid x, v_{< j_{r}})}.\tag{4}
$$

### Strict Cross-Tokenizer Objective

$$
\mathcal{L}_{1:1}(\theta) = \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_{\theta}(\cdot | x)} \left[ \sum_{r \in \mathcal{A}_{1:1}(y)} \mathrm{KL}\left(\bar{\pi}_{\theta}(\cdot | x, y_{< i_{r}}) \| \bar{\pi}_{\mathrm{T}}(\cdot | x, v_{< j_{r}})\right) \right].\tag{5}
$$

### Span Supervision (MSE on Mismatch Groups)

For mismatch groups, the probabilities of observed token paths are computed as:

$$
q_{\theta}^{(r)} = \prod_{i \in S_{r}^{\theta}} \pi_{\theta}(y_{i} \mid x, y_{< i}), \quad q_{\mathrm{T}}^{(r)} = \prod_{j \in S_{r}^{\mathrm{T}}} \pi_{\mathrm{T}}(v_{j} \mid x, v_{< j}).\tag{7}
$$

The span loss uses MSE on log-probabilities:

$$
\mathcal{L}_{\text{span}}(\theta) = \mathbb{E}_{x \sim \mathcal{D}, y \sim \pi_{\theta}(\cdot | x)} \left[ \sum_{r \in \mathcal{A}_{\text{mis}}(y)} \left(\log q_{\theta}^{(r)} - \log q_{\mathrm{T}}^{(r)}\right)^{2} \right].\tag{8}
$$

Total loss combines both terms:

$$
\mathcal{L}_{\lambda}(\theta) = \mathcal{L}_{1:1}(\theta) + \lambda \mathcal{L}_{\text{span}}(\theta).\tag{9}
$$

### Experimental Setup

- **Model pairs**: Qwen2.5-7B-Instruct→Llama-3.2-3B-Instruct, Granite-4.1-8B→Phi-4-mini-instruct, Granite-4.1-8B→Qwen2.5-7B-Base
- **Data**: 20,000 prompts (10,000 math from DAPO-Math-17K, 10,000 code from CodeForces)
- **Evaluation**: MATH500, GSM8K, AIME-2024/2025/2026, AMC23, Minerva-Math (math); HumanEval, MBPP, LiveCodeBench (code)
- **Metrics**: mean@32 for math, mean@8 for code

### Gradient Diagnostics

Directional agreement (cosine) and relative magnitude (norm ratio) between strict and span gradients:

$$
c_{\mathrm{mis}} = \frac{\langle g_{1:1}, g_{\mathrm{mis}}\rangle}{\|g_{1:1}\|_{2} \|g_{\mathrm{mis}}\|_{2}}, \quad \rho = \frac{\|g_{\mathrm{mis}}\|_{2}}{\|g_{1:1}\|_{2}}.\tag{13}
$$

---

## Empirical Validation / Results

### 1. Strict Alignment Coverage vs. Static Vocabulary Overlap

**Table 1**: Static vocabulary Jaccard overlap and strict token coverage on student trajectories

| Teacher→Student | $J_{vocab}$ | Coverage | Steps 1-100 | Steps 1-20 | Steps 41-60 | Steps 81-100 |
|---|---|---|---|---|---|---|
| Qwen→Llama | 64.32 | $C_{1:1}^{\theta}$ | 93.56 | 94.63 | 93.20 | 93.43 |
| | | $C_{1:1}^{\text{T}}$ | 85.91 | 88.65 | 84.90 | 85.36 |
| Granite→Phi | 39.49 | $C_{1:1}^{\theta}$ | 96.98 | 96.65 | 96.95 | 97.12 |
| | | $C_{1:1}^{\text{T}}$ | 97.26 | 95.83 | 97.50 | 97.96 |
| Granite→Qwen | 64.87 | $C_{1:1}^{\theta}$ | 85.57 | 85.23 | 85.06 | 85.88 |
| | | $C_{1:1}^{\text{T}}$ | 82.82 | 81.15 | 82.77 | 83.86 |

> **Key insight**: Granite→Phi has the *lowest* Jaccard overlap (39.49%) but the *highest* strict token coverage (>96%). Substantial static vocabulary mismatch can coexist with strict alignment at most positions on student-generated trajectories.

### 2. Span Supervision Reduces Accuracy

All three pairs attain their highest full average at $\lambda = 0$ (strict only). The 18 positive-weight settings score 0.27–1.20 percentage points below their respective strict baselines, with an overall downward trend as λ increases.

### 3. Probability Mass Concentrates on Shared Vocabulary

The shared vocabulary retains 99.69–99.90% of teacher mass and 98.99–99.81% of student mass on average. At k=16, the student-selected top-k subset retains at least 93.54% of teacher mass and 94.55% of student mass.

### 4. Top-k Subset Training Results

**Table 2**: Cross-tokenizer distillation accuracy (%) — Bold marks the best distilled score per column

| Method | Qwen→Llama Math | Qwen→Llama Code | Qwen→Llama Full | Granite→Phi Math | Granite→Phi Code | Granite→Phi Full | Granite→Qwen Math | Granite→Qwen Code | Granite→Qwen Full |
|---|---|---|---|---|---|---|---|---|---|
| Base | 18.32 | 35.60 | 26.96 | 33.06 | 40.04 | 36.55 | 24.61 | 34.51 | 29.56 |
| ULD | 20.51 | 37.74 | 29.12 | 35.24 | 42.01 | 38.63 | 33.36 | 20.87 | 27.12 |
| Extended ULD | 23.01 | 39.26 | 31.13 | 35.28 | 41.42 | 38.35 | 33.14 | 32.00 | 32.57 |
| GOLD | 21.21 | 33.10 | 27.16 | 36.22 | 47.29 | 41.75 | 38.26 | 54.84 | 46.55 |
| SimCT | 25.31 | 37.86 | 31.59 | 36.67 | 46.70 | 41.68 | 38.64 | 53.56 | 46.10 |
| **Strict full** | **26.60** | 39.11 | **32.86** | **37.49** | **47.49** | **42.49** | **39.85** | 54.84 | 47.35 |
| **Strict top-16** | 26.86 | 38.42 | 32.64 | 37.23 | 47.29 | 42.26 | 39.07 | 55.06 | 47.06 |
| **Strict top-128** | 26.31 | 38.44 | 32.38 | 37.27 | 47.82 | 42.54 | 39.67 | **55.62** | **47.65** |

Key findings:
- All three strict variants exceed all four alternatives in both math and full averages on every model pair
- Top-16 retains at least 96% of the full-average improvement from Base to strict full
- Top-16 is ahead of the strongest alternative in the full average by 0.51–1.05 percentage points
- No consistent further gain from increasing subset size beyond 16

### 5. Gradient Diagnostics

- **Directional agreement**: Mismatch cosines are near zero for Qwen→Llama and Granite→Phi; Granite→Qwen starts with negative agreement approaching zero later. The split-half strict cosine remains higher than the mismatch–strict cosine at every checkpoint.
- **Relative scale**: The span-to-strict gradient norm ratio increases across measured checkpoints in all three pairs, meaning a fixed positive weight on the span loss gives it increasing gradient norm relative to the strict component.

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **Coverage ≠ Learning Value**: The paper provides systematic evidence that expanding structural supervision coverage does not guarantee better distillation. This challenges the implicit assumption in prior work (e.g., SimCT) that recovering excluded supervision signals is beneficial.

2. **Support Concentration Under Heterogeneous Tokenizers**: The finding that a compact top-16 subset retains most distillation gains extends prior work on support concentration (Fu et al., 2026; Li et al., 2026) to the cross-tokenizer setting, showing that student-selected supports remain effective even when vocabularies only partially overlap.

3. **Gradient Interaction as Diagnostic Tool**: The gradient cosine and norm ratio analyses provide a mechanistic explanation for why span supervision hurts: weakly aligned gradients with growing relative magnitude can interfere with the strict objective.

### Practical Implications

1. **Efficient Cross-Tokenizer Distillation**: Practitioners can restrict distillation to strict 1:1 positions with a top-16 shared vocabulary subset, achieving near-full performance with substantially lower computational cost.

2. **Supervision Quality Assessment**: The paper suggests that before adding new supervision signals, practitioners should evaluate their directional agreement with existing losses and their relative gradient scale.

3. **Baseline Comparisons**: Strict full and Strict top-16 provide strong baselines that outperform more complex cross-tokenizer methods (ULD, GOLD, SimCT), suggesting that simpler approaches may be more effective.

---

## Conclusion

The paper re-examines the role of alignment coverage in Cross-Tokenizer OPD across three heterogeneous teacher–student pairs on mathematical reasoning and code generation. The main takeaways are:

1. **Strict alignment is already substantial**: Strict 1:1 groups cover most student-generated tokens despite large static vocabulary mismatch.

2. **Shared vocabulary retains most probability mass**: On student-generated responses, the shared vocabulary carries nearly all teacher and student predictive probability at strict positions.

3. **Compact supervision is effective**: Restricting reverse KL to a student-selected top-16 subset retains at least 96% of full shared-vocabulary OPD improvement, outperforming all evaluated cross-tokenizer baselines.

4. **Span supervision is harmful**: Adding span log-probability MSE achieves complete supervision coverage but reduces accuracy. Gradient diagnostics show span gradients have weak/negative directional agreement with strict gradients and grow in relative magnitude during training.

**Future directions** suggested by this work include:
- Developing principled methods for assessing supervision reliability before adding new training signals
- Investigating whether other forms of mismatch-group supervision (beyond MSE on log-probabilities) could be more effective
- Extending the analysis to other domains and larger model scales (the paper includes preliminary results with a 235B teacher in the appendix)

---

_Markdown view of https://picx.dev/p/tlUP8j, served by PicX — AI-generated visual whiteboard summaries of research papers._
