# RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

> RouterInterp shows MoE experts specialize in disjoint collections of unrelated micro-domains, not coherent semantic domains, challenging the standard specialisation hypothesis.

- **Source:** [arXiv](https://arxiv.org/abs/2610.11775)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/5j44Mp
- **Whiteboard:** https://picx.dev/p/5j44Mp/image

## Summary

# RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

## Summary (Overview)

- **Core contribution**: The paper introduces the **Superposed Specialisation Hypothesis (SSH)**, which posits that Mixture-of-Experts (MoE) models' experts specialize in *disjoint collections of unrelated micro-domains* rather than single coherent semantic domains, challenging the commonly assumed Domain Specialisation Hypothesis (DSH).

- **Novel method**: The authors propose **RouterInterp**, an SAE-based interpretability method that identifies sparse autoencoder features most predictive of routing decisions and generates unified natural language explanations of expert behavior.

- **Key empirical finding**: Expert routing is highly polysemantic—each expert's most predictive SAE latents span ~11 semantically distinct clusters (out of 20), far from the single cluster predicted by DSH.

- **Performance results**: RouterInterp achieves explanation scores of 0.49 on gpt-oss-20b and 0.60 on OLMoE-1B-7B, outperforming token-statistics-based methods by ~65% and prior expert-impact methods by ~28% on gpt-oss-20b.

- **Theoretical contribution**: The paper provides two arguments for why SSH holds: interference minimization (analogous to superposition theory) and load-balancing pressure during training.

## Introduction and Theoretical Foundation

### Background

Sparse Mixture-of-Experts (MoE) transformers scale more efficiently than dense models by routing tokens to modular expert networks, activating only 2–15% of parameters per input. The strong performance of MoE models has often been attributed to **expert specialisation**—the idea that each expert learns to handle a subset of the input distribution.

However, prior interpretability efforts have struggled to recover clear, interpretable patterns of expert specialisation. The paper identifies a crucial distinction between two forms of the Specialisation Hypothesis:

**Domain Specialisation Hypothesis (DSH)**: Experts specialize in different *semantically coherent domains* (e.g., medical text, mathematical reasoning, romance language translation). Semantically similar micro-domains cluster within an expert.

**Superposed Specialisation Hypothesis (SSH)**: Experts specialize in *disjoint collections of features* spanning multiple unrelated micro-domains. This mirrors the superposition phenomenon observed in neural networks.

### Theoretical Foundation

The key insight is the **Pigeonhole Principle argument**: real-world data contains far more fine-grained categories (micro-domains) than any MoE layer has experts. When $D > E$ (more micro-domains than experts), multiple micro-domains must inevitably be routed to the same expert. The question becomes whether those micro-domains are similar (DSH) or dissimilar (SSH).

### Motivation for SSH

**Argument 1 — Interference Minimisation**: Drawing on superposition theory (Elhage et al., 2022), the authors argue that semantically distinct micro-domains that rarely co-activate can share an expert with minimal interference noise, enabling "Computation in Superposition" (Hänni et al., 2024).

**Argument 2 — Load Balancing**: Load-balancing losses encourage uniform expert utilization. With small batch sizes, a batch contains few macro-domains, but the router must distribute tokens from these few domains across all experts—necessarily routing dissimilar inputs to the same expert.

## Methodology

### Sparse Autoencoders (SAEs)

The paper uses Top-K Sparse Autoencoders to decompose model activations into interpretable features. The SAE architecture maps an activation vector $\boldsymbol{x} \in \mathbb{R}^{N}$ to a sparse latent representation $\boldsymbol{z} \in \mathbb{R}^{F}$ (where $F > N$):

$$
\begin{array}{l} \boldsymbol {z} = \sigma (\boldsymbol {W} _ {\mathrm{enc}} (\boldsymbol {x} - \boldsymbol {b} _ {\mathrm{pre}}) + \boldsymbol {b} _ {\mathrm{enc}}) \\ \hat {\boldsymbol {x}} = \boldsymbol {W} _ {\mathrm{dec}} \boldsymbol {z} + \boldsymbol {b} _ {\mathrm{pre}} \end{array}
$$

where $\boldsymbol{W}_{\mathrm{enc}} \in \mathbb{R}^{F \times N}$, $\boldsymbol{W}_{\mathrm{dec}} \in \mathbb{R}^{N \times F}$, and $\sigma = \mathrm{TopK}(\cdot, s)$ retains only the $s$ largest activations.

### MoE Routing

In a standard MoE layer, a router network computes routing logits and converts them to a probability distribution over experts via softmax:

$$
g _ {i} (\boldsymbol {x}) = \frac {\exp (h (\boldsymbol {x}) _ {i})}{\sum_ {j = 1} ^ {E} \exp (h (\boldsymbol {x}) _ {j})}
$$

Only the top-$k$ experts are selected, and the layer output is:

$$
\boldsymbol {y} = \sum_ {i \in \mathcal {T}} g _ {i} (\boldsymbol {x}) \cdot E _ {i} (\boldsymbol {x})
$$

### RouterInterp Method

RouterInterp operates in three stages:

1. **Identifying Features**: For each expert $E_i$, select the top-$n$ SAE features most predictive of routing by approximating the effect of ablating each feature on expert selection. Each feature is assigned to at most one expert.

2. **Collecting Activations**: For each selected feature $f \in \mathbb{F}_i$, collect context windows where $f$ is active, forming:
   - **Positive set** $\mathcal{P}_{i,f}$: $f$ active AND token routed to $E_i$
   - **Negative set** $\mathcal{N}_{i,f}$: $f$ active but token NOT routed to $E_i$

3. **Explaining**: Prompt a language model with positive-negative contrast groups per feature to synthesize a unified natural language description of when expert $E_i$ is selected.

### Evaluation Setup

- **Models**: OLMoE-1B-7B (64 experts, $k=8$) and gpt-oss-20b (32 experts, $k=4$)
- **SAEs**: Top-K SAEs trained on 100M tokens (OLMoE) with 32,768 features, sparsity $s=32$; BatchTopK SAEs for gpt-oss with 131,072 features, sparsity $s \in \{64, 128\}$
- **Scoring**: Adapted AutoInterp Detection scoring—a scorer LLM uses the explanation to predict expert activation on held-out examples, with F1 as the metric
- **Baselines**: Unigram Lookup (token co-occurrence), Expert Impact AutoInterp (Herbst et al., 2026)

## Empirical Validation / Results

### Evidence for SSH

The paper tests whether each expert's top-20 predictive SAE latents are mutually similar (DSH) or dissimilar (SSH). They define $G(E_i)$ as the number of semantically distinct clusters that expert $E_i$'s top-$n$ latents fall into:

- **DSH predicts**: $G(E_i) = 1$ (single coherent domain)
- **SSH predicts**: $G(E_i) \approx n$ (each latent in its own cluster)
- **Observed**: Mean $G$ = 11.3 on gpt-oss-20b (layer 20) and 10.8 on OLMoE-1B-7B (layer 15), out of upper bound $n=20$

The observed values lie far above the DSH prediction of $G=1$, providing strong evidence for SSH.

### RouterInterp Performance

**Figure 4 results** (mean F1 scores):

| Method | gpt-oss-20b | OLMoE-1B-7B |
|--------|------------|-------------|
| RouterInterp | **0.492** (s=128) | **0.602** |
| Expert Impact AutoInterp | 0.383 | 0.305 |
| Unigram Lookup | 0.299 | 0.342 |

### Ablation Study (Table 1, gpt-oss-20b)

| Method | Layer 4 | Layer 8 | Layer 12 | Layer 16 | Layer 20 |
|--------|---------|---------|----------|----------|----------|
| Unigram Lookup | 0.295 | 0.300 | 0.278 | 0.305 | 0.315 |
| Bigram Lookup | 0.309 | 0.320 | 0.283 | 0.320 | 0.331 |
| Unigram AutoInterp | 0.332 | 0.406 | 0.320 | 0.356 | 0.348 |
| Bigram AutoInterp | 0.429 | 0.453 | 0.385 | 0.407 | 0.411 |
| Expert Activations AutoInterp | 0.415 | 0.475 | 0.345 | 0.425 | 0.425 |
| **RouterInterp (s=64)** | **0.490** | **0.542** | **0.465** | **0.509** | **0.470** |
| **RouterInterp (s=128)** | **0.531** | **0.509** | **0.456** | **0.516** | **0.448** |

### Key Ablation Insights

- Verbalizing token statistics with an LLM improves explanations only modestly (0.30→0.35 mean F1)
- Replacing token lists with full activating passages does not help (Expert Activations AutoInterp ≈ Bigram AutoInterp at 0.417 mean F1)
- Bigram descriptions are **narrow** (higher precision 0.39, lower recall 0.59); full-window descriptions are **broad** (lower precision 0.30, higher recall 0.81)
- Only when evidence is grouped by SAE feature do precision and recall improve together (P 0.47 / R 0.64; 0.495 mean F1, s=64)

## Theoretical and Practical Implications

### Implications for Interpretability

1. **Explains prior failures**: The SSH explains why previous work assuming monosemantic experts (implicitly following DSH) was unable to successfully interpret routing behavior.

2. **Reconciles apparent specialisation**: Cases of apparent domain specialisation (e.g., code generation experts, language-specific experts, safety refusal experts) are consistent with SSH—an expert may participate in a task without that being its only function.

3. **SAE-based discovery**: The success of RouterInterp supports the view that SAEs are useful for "discovery of unknowns" rather than acting on known concepts (Peng et al., 2026).

### Theoretical Implications

1. **Computation in Superposition**: Experts may group seemingly unrelated inputs that require a set of disjoint computational operations, suggesting a **functional similarity** (rather than input similarity) basis for routing.

2. **Design implications**: Understanding whether experts cluster by input similarity or functional similarity could inform the design of more efficient and interpretable MoE architectures.

3. **Scaling insights**: With much larger numbers of experts, routing becomes increasingly monosemantic and interpretable (Park et al., 2025), suggesting a trade-off between interpretability and compute efficiency.

### Practical Implications

- RouterInterp provides a scalable method for generating accurate explanations of expert routing, potentially enabling:
  - Better debugging and auditing of MoE models
  - More targeted interventions in expert behavior
  - Improved transparency for safety-critical applications of foundation models

## Conclusion

The paper makes three main contributions:

1. **Theoretical**: Introduces and provides evidence for the Superposed Specialisation Hypothesis, formalizing the distinction between domain-level and feature-level specialisation in MoE models.

2. **Empirical**: Demonstrates that expert routing is highly polysemantic, with experts' predictive features spanning many semantically distinct clusters.

3. **Methodological**: Presents RouterInterp, which achieves state-of-the-art explanation scores (0.49 on gpt-oss-20b, 0.60 on OLMoE-1B-7B) by leveraging SAE features to enumerate experts' disjoint micro-domains.

### Future Directions

- **Mechanisms of specialisation**: Why do experts converge to particular feature combinations? Does load-balancing force redundancy?
- **Functional vs. input similarity**: Do co-routed tokens share semantic content or require similar downstream transformations?
- **Developmental interpretability**: How does routing evolve across training checkpoints?
- **Scaling**: Apply RouterInterp to frontier architectures with more parameters, shared experts, and different routing methods (e.g., Expert Choice).

### Limitations

- Dependence on SAE quality (performance correlates with reconstruction quality/FVU)
- Explanations grow in size with more features; compression and interactive formats (e.g., Neuronpedia dashboards) could help
- Evaluated only on 1B–20B parameter models; scaling to frontier architectures remains open

---

_Markdown view of https://picx.dev/p/5j44Mp, served by PicX — AI-generated visual whiteboard summaries of research papers._
