Full text not available for this paper

Summary (Overview)

  • Central claim: Routing drift alone is insufficient evidence of routing failure in merged Mixture-of-Experts (MoE) LLMs; source-informed corrections must be judged by their task-level intervention effects, not by structural routing disagreement.
  • Key finding 1: Across DeepSeekMoE, OLMoE, and Qwen3-MoE under Average and Task Arithmetic merging, 77.7–96.9% of changed-route events are attributable to input (representation) shifts rather than router parameter changes.
  • Key finding 2: Structural routing differences (e.g., JS divergence) poorly predict next-token likelihood gains from source-route restoration (AUROC 0.47–0.52, near chance), and different expert selections can produce directionally similar mixture outputs (cosine similarity 0.896–0.976).
  • Key finding 3: The paper formalizes routing failure as intervention-relative recoverable task loss with non-routing parameters fixed, and demonstrates that controlled router corruption is recoverable (12.50–32.81 pp) while source-route restoration yields no reliable task benefit.
  • Key finding 4: The proposed Selective Router Repair (SRR) method—which fits likelihood-weighted source-derived expert-pair corrections—shows no reliable evidence that source-likelihood advantages identify beneficial corrections or improve average task performance (−0.056 to +0.139 pp across settings).

Introduction and Theoretical Foundation

Background: Model merging combines specialized LLMs without joint retraining. For MoE models, merging changes not only parameters but also token-to-expert routing—termed routing drift. Recent methods (e.g., HARC) treat routing mismatch as "routing breakdown" and realign merged routers.

Core distinction: The paper distinguishes two concepts:

  • Routing drift: what changed in expert assignment and routing probabilities
  • Routing failure: degradation in task-relevant behavior attributable to routing

Theoretical basis: A token's route depends jointly on (1) its router-input representation (hidden state), (2) the router parameters, and (3) the weighted outputs of selected experts. Therefore, routing drift alone reveals neither why routing changed nor whether it caused harm. The MoE layer output is:

o(ℓ)(h;S,g)=∑e∈SgeEe(ℓ)(h)o^{(\ell)}(h; S, g) = \sum_{e \in S} g_e E_e^{(\ell)}(h)

where SS is the selected expert set, geg_e the routing weight of expert ee, and Ee(ℓ)(h)E_e^{(\ell)}(h) its output at layer ℓ\ell.

Research questions: (1) Does routing drift after MoE merging actually indicate routing failure? (2) What evidence should justify repair?

Methodology

Models and merging methods: DeepSeekMoE-16B-A3B (Top-6 routing, 27 sparse layers), OLMoE-7B-A1B (Top-8, 16 layers), Qwen3-30B-A3B (Top-8, 48 layers), each evaluated under Average, Task Arithmetic (TA), TIES, and WUDI-Merge.

Crossed intervention analysis: For source and merged quantities, four states are evaluated:

rSS=R(WS,hS),rMS=R(WM,hS),rSM=R(WS,hM),rMM=R(WM,hM)r_{SS} = R(W_S, h_S), \quad r_{MS} = R(W_M, h_S), \quad r_{SM} = R(W_S, h_M), \quad r_{MM} = R(W_M, h_M)

where R(W,h)=Top-k(Wh)R(W, h) = \text{Top-}k(Wh) is the expert selection, WW denotes router parameters, and hh the router input. Conditioning on rSS≠rMMr_{SS} \neq r_{MM}: a change is representation-induced when rSM≠rSSr_{SM} \neq r_{SS} but rMS=rSSr_{MS} = r_{SS}; router-parameter-induced is the converse.

Task-grounded diagnosis: Define intervention-relative recoverable loss:

HD(π;π′)=E(x,y)∼D[ℓtask(Mϕ,π;x,y)−ℓtask(Mϕ,π′;x,y)]H_D(\pi; \pi') = \mathbb{E}_{(x,y) \sim D}[\ell_{task}(M_{\phi,\pi}; x, y) - \ell_{task}(M_{\phi,\pi'}; x, y)]

where π\pi and π′\pi' are baseline and alternative routing policies, ϕ\phi are fixed non-routing parameters. Positive values indicate recoverable loss.

Selective Router Repair (SRR): Constructs source–base preference profiles per prompt:

uˉi=∑t∈iωt center(ztS−ztB)∑t∈iωt,ωt=σ((ntB−ntS)/τ)\bar{u}_i = \frac{\sum_{t \in i} \omega_t \, \text{center}(z^S_t - z^B_t)}{\sum_{t \in i} \omega_t}, \quad \omega_t = \sigma((n^B_t - n^S_t)/\tau)

where zz are router logits, nn are NLLs, σ\sigma is sigmoid, and center subtracts the expert-wise mean. For each selected pair j=(aj,bj)j = (a_j, b_j), the fitting objective is:

vj∗=arg⁡min⁡v12∑t=1Nρtj((htM)⊤v−r~tj)2+λj2∥v∥22v^*_j = \arg\min_v \frac{1}{2} \sum_{t=1}^{N} \rho_{tj} \left( (h^M_t)^\top v - \tilde{r}_{tj} \right)^2 + \frac{\lambda_j}{2} \|v\|_2^2

with λj=γmax⁡(N−1∑tρtj∥htM∥22,10−12)\lambda_j = \gamma \max\left( N^{-1} \sum_t \rho_{tj} \|h^M_t\|_2^2, 10^{-12} \right). Router rows update as waj′=wajM+η2vj∗w'_{a_j} = w^M_{a_j} + \frac{\eta}{2} v^*_j and wbj′=wbjM−η2vj∗w'_{b_j} = w^M_{b_j} - \frac{\eta}{2} v^*_j.

Evaluation: 8 benchmarks (MMLU, HellaSwag, ARC-C, ARC-E, PIQA, WinoGrande, BoolQ, GSM8K) with 35,326 matched items per pair, five evaluation runs, paired item-bootstrap 95% confidence intervals.

Empirical Validation / Results

Routing origin attribution (Table A2, Figure 2):

Architecture / ParentChanged route (%)Representation (%)Router parameters (%)
DeepSeekMoE / Average33.9 [32.2, 35.7]77.7 [77.1, 78.2]1.79 [1.65, 1.94]
DeepSeekMoE / TA46.7 [44.9, 48.7]78.5 [78.1, 79.0]1.14 [1.04, 1.23]
OLMoE / Average42.7 [41.3, 44.1]85.8 [85.4, 86.1]0.62 [0.57, 0.68]
OLMoE / TA54.3 [52.9, 55.7]87.0 [86.6, 87.4]0.36 [0.33, 0.40]
Qwen3-MoE / Average26.4696.870.034
Qwen3-MoE / TA33.2096.900.040

Structural predictors near chance: JS divergence predicts source-route NLL gain with AUROC 0.47–0.52 (chance = 0.5). None of the 16 primary tests survives Holm correction.

Mixture output similarity (Figure 4): Source-route and native merged-route mixtures have mean cosine 0.896–0.976, while expert-pair maximum cosine is only 0.041–0.096. Observed routes exceed matched-random controls by 0.112–0.147 in cosine.

Controlled corruption recovery (Figure 5): Permuting router logits in OLMoE layers recovers 12.50–14.84 pp (4 layers) and 26.56–32.81 pp (16 layers) via clean-route replay. All eight comparisons at k=4, 16 pass Holm correction.

Natural routing alternatives (Table 1): Source-route replay and LC replay yield no strictly positive intervals for accuracy or margin changes across all settings.

SRR task results (Table 2): Average score changes range from −0.056 to +0.139 pp across eight OLMoE/Qwen3-MoE settings. All 95% CIs include zero. HARC and SRR each lead in four of eight settings.

Local utility diagnostics: Source and fitted directions agree in sign on 83.4% of supported events, yet neither pooled UU (gain over native) nor DD (directional gain) establishes improvement. Matched selection effects show no enrichment (Table 3):

ContrastEstimate95% CI
Native (UU)−0.341[−3.323, 2.631]
Opposite (DD)−0.997[−5.653, 3.868]

Theoretical and Practical Implications

Theoretical implications:

  • Routing drift is a structural phenomenon; routing failure is a causal, task-relative phenomenon. These must not be conflated.
  • The paper provides a formal framework (intervention-relative recoverable loss) for diagnosing routing failure that separates candidate construction from demonstrated recovery.
  • Different routes can preserve similar mixture outputs (functional redundancy), explaining why routing disagreement need not imply harm.

Practical implications:

  • Post-merge routing "repair" methods (e.g., HARC) should be evaluated against task-level intervention effects, not source-route agreement.
  • Source-likelihood advantages do not reliably identify beneficial router corrections; repair decisions require task-grounded validation.
  • The released analysis toolkit enables controlled counterfactual interventions and paired token-/task-level evaluation for future MoE merging research.

Conclusion

The paper demonstrates that routing drift alone is insufficient evidence of routing failure in merged MoE LLMs. Most post-merge route changes are representation-induced rather than router-parameter-induced, structural differences poorly predict intervention gains, and different routes can preserve similar mixture outputs. The formalization of routing failure as intervention-relative recoverable task loss—with non-routing parameters fixed—provides a principled diagnostic framework. Controlled corruption tests confirm recoverability under deliberate router corruption, while source-route restoration shows no reliable task benefit. The proposed SRR case study finds no evidence that source-likelihood advantages identify beneficial corrections or that fitted updates improve average task performance.

Future directions: (1) Exploring alternative routing interventions beyond source-route restoration that might yield recoverable task loss; (2) developing selection criteria for router repair grounded in task-level intervention effects rather than source agreement; (3) extending the analysis to additional MoE architectures and merging methods; (4) investigating whether other supervision signals (beyond source-likelihood) can identify beneficial local corrections.

Related papers