RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

Summary (Overview)

  • Core contribution: The paper introduces the Superposed Specialisation Hypothesis (SSH), which posits that Mixture-of-Experts (MoE) models' experts specialize in disjoint collections of unrelated micro-domains rather than single coherent semantic domains, challenging the commonly assumed Domain Specialisation Hypothesis (DSH).

  • Novel method: The authors propose RouterInterp, an SAE-based interpretability method that identifies sparse autoencoder features most predictive of routing decisions and generates unified natural language explanations of expert behavior.

  • Key empirical finding: Expert routing is highly polysemantic—each expert's most predictive SAE latents span ~11 semantically distinct clusters (out of 20), far from the single cluster predicted by DSH.

  • Performance results: RouterInterp achieves explanation scores of 0.49 on gpt-oss-20b and 0.60 on OLMoE-1B-7B, outperforming token-statistics-based methods by ~65% and prior expert-impact methods by ~28% on gpt-oss-20b.

  • Theoretical contribution: The paper provides two arguments for why SSH holds: interference minimization (analogous to superposition theory) and load-balancing pressure during training.

Introduction and Theoretical Foundation

Background

Sparse Mixture-of-Experts (MoE) transformers scale more efficiently than dense models by routing tokens to modular expert networks, activating only 2–15% of parameters per input. The strong performance of MoE models has often been attributed to expert specialisation—the idea that each expert learns to handle a subset of the input distribution.

However, prior interpretability efforts have struggled to recover clear, interpretable patterns of expert specialisation. The paper identifies a crucial distinction between two forms of the Specialisation Hypothesis:

Domain Specialisation Hypothesis (DSH): Experts specialize in different semantically coherent domains (e.g., medical text, mathematical reasoning, romance language translation). Semantically similar micro-domains cluster within an expert.

Superposed Specialisation Hypothesis (SSH): Experts specialize in disjoint collections of features spanning multiple unrelated micro-domains. This mirrors the superposition phenomenon observed in neural networks.

Theoretical Foundation

The key insight is the Pigeonhole Principle argument: real-world data contains far more fine-grained categories (micro-domains) than any MoE layer has experts. When D>ED > E (more micro-domains than experts), multiple micro-domains must inevitably be routed to the same expert. The question becomes whether those micro-domains are similar (DSH) or dissimilar (SSH).

Motivation for SSH

Argument 1 — Interference Minimisation: Drawing on superposition theory (Elhage et al., 2022), the authors argue that semantically distinct micro-domains that rarely co-activate can share an expert with minimal interference noise, enabling "Computation in Superposition" (Hänni et al., 2024).

Argument 2 — Load Balancing: Load-balancing losses encourage uniform expert utilization. With small batch sizes, a batch contains few macro-domains, but the router must distribute tokens from these few domains across all experts—necessarily routing dissimilar inputs to the same expert.

Methodology

Sparse Autoencoders (SAEs)

The paper uses Top-K Sparse Autoencoders to decompose model activations into interpretable features. The SAE architecture maps an activation vector x∈RN\boldsymbol{x} \in \mathbb{R}^{N} to a sparse latent representation z∈RF\boldsymbol{z} \in \mathbb{R}^{F} (where F>NF > N):

z=σ(Wenc(x−bpre)+benc)x^=Wdecz+bpre\begin{array}{l} \boldsymbol {z} = \sigma (\boldsymbol {W} _ {\mathrm{enc}} (\boldsymbol {x} - \boldsymbol {b} _ {\mathrm{pre}}) + \boldsymbol {b} _ {\mathrm{enc}}) \\ \hat {\boldsymbol {x}} = \boldsymbol {W} _ {\mathrm{dec}} \boldsymbol {z} + \boldsymbol {b} _ {\mathrm{pre}} \end{array}

where Wenc∈RF×N\boldsymbol{W}_{\mathrm{enc}} \in \mathbb{R}^{F \times N}, Wdec∈RN×F\boldsymbol{W}_{\mathrm{dec}} \in \mathbb{R}^{N \times F}, and σ=TopK(⋅,s)\sigma = \mathrm{TopK}(\cdot, s) retains only the ss largest activations.

MoE Routing

In a standard MoE layer, a router network computes routing logits and converts them to a probability distribution over experts via softmax:

gi(x)=exp⁡(h(x)i)∑j=1Eexp⁡(h(x)j)g _ {i} (\boldsymbol {x}) = \frac {\exp (h (\boldsymbol {x}) _ {i})}{\sum_ {j = 1} ^ {E} \exp (h (\boldsymbol {x}) _ {j})}

Only the top-kk experts are selected, and the layer output is:

y=∑i∈Tgi(x)⋅Ei(x)\boldsymbol {y} = \sum_ {i \in \mathcal {T}} g _ {i} (\boldsymbol {x}) \cdot E _ {i} (\boldsymbol {x})

RouterInterp Method

RouterInterp operates in three stages:

  1. Identifying Features: For each expert EiE_i, select the top-nn SAE features most predictive of routing by approximating the effect of ablating each feature on expert selection. Each feature is assigned to at most one expert.

  2. Collecting Activations: For each selected feature f∈Fif \in \mathbb{F}_i, collect context windows where ff is active, forming:

    • Positive set Pi,f\mathcal{P}_{i,f}: ff active AND token routed to EiE_i
    • Negative set Ni,f\mathcal{N}_{i,f}: ff active but token NOT routed to EiE_i
  3. Explaining: Prompt a language model with positive-negative contrast groups per feature to synthesize a unified natural language description of when expert EiE_i is selected.

Evaluation Setup

  • Models: OLMoE-1B-7B (64 experts, k=8k=8) and gpt-oss-20b (32 experts, k=4k=4)
  • SAEs: Top-K SAEs trained on 100M tokens (OLMoE) with 32,768 features, sparsity s=32s=32; BatchTopK SAEs for gpt-oss with 131,072 features, sparsity s∈{64,128}s \in \{64, 128\}
  • Scoring: Adapted AutoInterp Detection scoring—a scorer LLM uses the explanation to predict expert activation on held-out examples, with F1 as the metric
  • Baselines: Unigram Lookup (token co-occurrence), Expert Impact AutoInterp (Herbst et al., 2026)

Empirical Validation / Results

Evidence for SSH

The paper tests whether each expert's top-20 predictive SAE latents are mutually similar (DSH) or dissimilar (SSH). They define G(Ei)G(E_i) as the number of semantically distinct clusters that expert EiE_i's top-nn latents fall into:

  • DSH predicts: G(Ei)=1G(E_i) = 1 (single coherent domain)
  • SSH predicts: G(Ei)≈nG(E_i) \approx n (each latent in its own cluster)
  • Observed: Mean GG = 11.3 on gpt-oss-20b (layer 20) and 10.8 on OLMoE-1B-7B (layer 15), out of upper bound n=20n=20

The observed values lie far above the DSH prediction of G=1G=1, providing strong evidence for SSH.

RouterInterp Performance

Figure 4 results (mean F1 scores):

Methodgpt-oss-20bOLMoE-1B-7B
RouterInterp0.492 (s=128)0.602
Expert Impact AutoInterp0.3830.305
Unigram Lookup0.2990.342

Ablation Study (Table 1, gpt-oss-20b)

MethodLayer 4Layer 8Layer 12Layer 16Layer 20
Unigram Lookup0.2950.3000.2780.3050.315
Bigram Lookup0.3090.3200.2830.3200.331
Unigram AutoInterp0.3320.4060.3200.3560.348
Bigram AutoInterp0.4290.4530.3850.4070.411
Expert Activations AutoInterp0.4150.4750.3450.4250.425
RouterInterp (s=64)0.4900.5420.4650.5090.470
RouterInterp (s=128)0.5310.5090.4560.5160.448

Key Ablation Insights

  • Verbalizing token statistics with an LLM improves explanations only modestly (0.30→0.35 mean F1)
  • Replacing token lists with full activating passages does not help (Expert Activations AutoInterp ≈ Bigram AutoInterp at 0.417 mean F1)
  • Bigram descriptions are narrow (higher precision 0.39, lower recall 0.59); full-window descriptions are broad (lower precision 0.30, higher recall 0.81)
  • Only when evidence is grouped by SAE feature do precision and recall improve together (P 0.47 / R 0.64; 0.495 mean F1, s=64)

Theoretical and Practical Implications

Implications for Interpretability

  1. Explains prior failures: The SSH explains why previous work assuming monosemantic experts (implicitly following DSH) was unable to successfully interpret routing behavior.

  2. Reconciles apparent specialisation: Cases of apparent domain specialisation (e.g., code generation experts, language-specific experts, safety refusal experts) are consistent with SSH—an expert may participate in a task without that being its only function.

  3. SAE-based discovery: The success of RouterInterp supports the view that SAEs are useful for "discovery of unknowns" rather than acting on known concepts (Peng et al., 2026).

Theoretical Implications

  1. Computation in Superposition: Experts may group seemingly unrelated inputs that require a set of disjoint computational operations, suggesting a functional similarity (rather than input similarity) basis for routing.

  2. Design implications: Understanding whether experts cluster by input similarity or functional similarity could inform the design of more efficient and interpretable MoE architectures.

  3. Scaling insights: With much larger numbers of experts, routing becomes increasingly monosemantic and interpretable (Park et al., 2025), suggesting a trade-off between interpretability and compute efficiency.

Practical Implications

  • RouterInterp provides a scalable method for generating accurate explanations of expert routing, potentially enabling:
    • Better debugging and auditing of MoE models
    • More targeted interventions in expert behavior
    • Improved transparency for safety-critical applications of foundation models

Conclusion

The paper makes three main contributions:

  1. Theoretical: Introduces and provides evidence for the Superposed Specialisation Hypothesis, formalizing the distinction between domain-level and feature-level specialisation in MoE models.

  2. Empirical: Demonstrates that expert routing is highly polysemantic, with experts' predictive features spanning many semantically distinct clusters.

  3. Methodological: Presents RouterInterp, which achieves state-of-the-art explanation scores (0.49 on gpt-oss-20b, 0.60 on OLMoE-1B-7B) by leveraging SAE features to enumerate experts' disjoint micro-domains.

Future Directions

  • Mechanisms of specialisation: Why do experts converge to particular feature combinations? Does load-balancing force redundancy?
  • Functional vs. input similarity: Do co-routed tokens share semantic content or require similar downstream transformations?
  • Developmental interpretability: How does routing evolve across training checkpoints?
  • Scaling: Apply RouterInterp to frontier architectures with more parameters, shared experts, and different routing methods (e.g., Expert Choice).

Limitations

  • Dependence on SAE quality (performance correlates with reconstruction quality/FVU)
  • Explanations grow in size with more features; compression and interactive formats (e.g., Neuronpedia dashboards) could help
  • Evaluated only on 1B–20B parameter models; scaling to frontier architectures remains open

Related papers