# DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition

> DivMoE achieves fine-grained MoE upcycling by combining domain-specialized experts with diversity-constrained routing, preventing routing collapse and outperforming all baselines across 15 benchmarks.

- **Source:** [arXiv](https://arxiv.org/abs/2610.11317)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/ZsznsA
- **Whiteboard:** https://picx.dev/p/ZsznsA/image

## Summary

# DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition

## Summary (Overview)

- **Key Problem Identified**: The paper uncovers a **fine-grained upcycling collapse pathology**—when fine-grained Mixture-of-Experts (MoE) experts are derived from a single source dense model, naive routing collapses and downstream accuracy drops to near-random levels (e.g., 23.2% average accuracy on Qwen3-1.7B, essentially matching from-scratch training at 22.2%).

- **Proposed Framework**: DIVMOE is the **first framework achieving fine-grained MoE Upcycling** with structurally-balanced routing, combining two innovations: (1) domain-specialized fine-grained expert initialization from multiple domain-adapted dense models, and (2) diversity-constrained routing that structurally guarantees each token activates experts from distinct domain groups.

- **Performance Gains**: DIVMOE consistently outperforms six upcycling baselines across two base models and 15 benchmarks (55.6% vs. 51.6% for the strongest baseline on Qwen3-1.7B), and strictly improves over the dense base model on every benchmark after Stage 2 continual pre-training.

- **Scalability Demonstration**: After supervised fine-tuning, a 12B-parameter DIVMOE model matches Moonlight-MoE (16B) at 64.5% average accuracy while outperforming a controlled NVIDIA-Upcycling baseline by 6.1 percentage points.

- **Counter-Intuitive Finding**: The structural diversity constraint *improves* performance on every domain-specific benchmark (HumanEval +1.8%, GSM8K +2.1%, MATH +1.4%, MBPP +2.3%), refuting the intuition that "coding tokens should only use coding experts."

## Introduction and Theoretical Foundation

### Background and Motivation

Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, enabling substantial capacity expansion while maintaining computational efficiency through sparse activation. Recent advances demonstrate the effectiveness of **fine-grained MoE designs**, where each feed-forward network (FFN) layer is partitioned into many small experts rather than a few large ones, contributing to the strong performance of models such as DeepSeek-V2 and Mixtral.

Training such models from scratch demands substantial compute, motivating **sparse upcycling**—converting pre-trained dense checkpoints into MoE architectures while inheriting their learned representations. However, the authors identify that fine-grained upcycling is fundamentally fragile when experts are derived from a single dense source.

### The Collapse Pathology

The paper's central observation is that no prior method simultaneously occupies three design axes that are jointly necessary:

1. **(a) Domain-specialized expert initialization**—experts derived from domain-adapted dense models
2. **(b) Fine-grained partitioning**—FFN layers split into many small experts
3. **(c) A structural mechanism preventing same-source co-activation**—hard constraints on routing

Existing methods fail on at least one axis:
- **Branch Train-Mix (BTX)** addresses (a) but uses coarse experts and standard top-k routing
- **NVIDIA Upcycling** addresses (b) but yields homogeneous expert initial scores that suppress routing diversity
- **Drop-Upcycling** targets diversity through randomness, sacrificing (a)

### Theoretical Foundation

Given a dense transformer with SwiGLU FFN:

$$\mathrm{FFN}(\mathbf{x}) = \mathbf{W}_{\mathrm{down}} \cdot \sigma(\mathbf{W}_{\mathrm{gate}}\mathbf{x}) \odot (\mathbf{W}_{\mathrm{up}}\mathbf{x})$$

where $\mathbf{W}_{\mathrm{gate}}, \mathbf{W}_{\mathrm{up}} \in \mathbb{R}^{d_{\mathrm{ff}} \times d}$ and $\mathbf{W}_{\mathrm{down}} \in \mathbb{R}^{d \times d_{\mathrm{ff}}}$, the model converts each FFN into an MoE layer. With $n = 4$ domain specialists, $m = 2$ shards per specialist, and $k = 2$ active experts, DIVMOE doubles total capacity while activating half the dense FFN parameters.

## Methodology

### 3.1 Domain-Specialized Fine-Grained Expert Initialization

**Domain specialists**: Starting from base model $\mathcal{M}_{\mathrm{base}}$, the method applies domain-adaptive continual pre-training (next-token-prediction loss on full sequences) on public domain-specific subsets to obtain $n$ specialists: $\{\mathcal{M}_{\mathrm{math}}, \mathcal{M}_{\mathrm{code}}, \mathcal{M}_{\mathrm{science}}, \mathcal{M}_{\mathrm{common}}\}$. Each specialist encodes distinct domain knowledge while retaining general capabilities.

**Expert construction**: Each specialist's FFN weights are vertically sliced into $m$ shards:

$$\mathbf{W}_{\mathrm{gate},i} \rightarrow \{\mathbf{W}_{\mathrm{gate},i}^{(1)}, \ldots, \mathbf{W}_{\mathrm{gate},i}^{(m)}\}$$

$$\mathbf{W}_{\mathrm{up},i} \rightarrow \{\mathbf{W}_{\mathrm{up},i}^{(1)}, \ldots, \mathbf{W}_{\mathrm{up},i}^{(m)}\}$$

$$\mathbf{W}_{\mathrm{down},i} \rightarrow \{\mathbf{W}_{\mathrm{down},i}^{(1)}, \ldots, \mathbf{W}_{\mathrm{down},i}^{(m)}\}$$

where expert $(i, j)$ has dimensions $(d_{\mathrm{ff}}/m) \times d$ for gate/up and $d \times (d_{\mathrm{ff}}/m)$ for down projections. Indices $i \in \{1, \ldots, n\}$ and $j \in \{1, \dots, m\}$ identify the specialist (domain group) and shard, respectively.

**Dense-equivalent initialization**: Following He et al., the method ensures stable initialization via virtual groups. Decomposing each specialist as $\mathbf{W}_i = \mathbf{W}_{\mathrm{base}} + \mathbf{\Delta}_i$ where $\mathbf{\Delta}_i$ captures domain adaptations:

$$\sum_{j=1}^{m} \mathrm{Expert}_{i,j}^{(j)}(\mathbf{x}) \approx \mathrm{FFN}_{\text{base}}(\mathbf{x})$$

This preserves the base model's output distribution at initialization while introducing structured diversity through the $\mathbf{\Delta}_i$ terms.

**Magnitude correction**: Expert weights are scaled by $\gamma$ to compensate for softmax attenuation of router scores:

$$\mathbf{W}_{\text{expert}} = \gamma \cdot \mathbf{W}_{\text{slice}}, \quad \gamma = \sqrt[3]{\frac{E \cdot G^2}{k}}$$

where $E$ is the expansion factor, $G$ is granularity, and $k$ is the active-expert count.

**Non-expert components**: Embeddings, attention, and normalization layers are initialized as the mean of the base model and all specialists:

$$\mathbf{W}_{\text{shared}} = \frac{1}{n+1}\left(\mathbf{W}_{\text{base}} + \sum_{i=1}^{n} \mathbf{W}_i\right)$$

### 3.2 Diversity-Constrained Routing

Standard MoE routing computes scores $\mathbf{s} = \mathbf{W}_r \mathbf{x} \in \mathbb{R}^N$ and selects $S = \mathrm{TopK}(\mathrm{Softmax}(\mathbf{s}), k)$. When experts are partitioned into $n$ domain groups $\{\mathcal{D}_1, \ldots, \mathcal{D}_n\}$, this top-k rule admits within-group concentration.

The **diversity constraint** imposes that each token activates at most one expert per group:

$$\forall i \in \{1, \dots, n\}: \quad |\mathcal{S} \cap \mathcal{D}_i| \leq 1$$

combined with $k \leq n$, guaranteeing that the $k$ selected experts originate from $k$ distinct domain groups.

**Algorithm 1: Diversity-Constrained Routing**

```
Input: hidden state x; router weights W_r; #groups n; experts/group m; active k
Output: selected experts S, weights w

s ← W_r x                          ▷ Router scores ∈ R^{n×m}

// Stage 1: Intra-group selection
for i = 1 to n do
    j_i* ← argmax_{j ∈ [m]} s_{i,j}   ▷ Best expert in group i
    s_i* ← s_{i,j_i*}                 ▷ Representative score
end for

// Stage 2: Inter-group selection
G ← TopK({s_1*, ..., s_n*}, k)        ▷ Top-k groups
S ← {(i, j_i*) : i ∈ G}               ▷ Selected experts

// Compute normalized routing weights
for i ∈ G do
    w_i ← exp(s_i*) / Σ_{j∈G} exp(s_j*)
end for
Return S, w
```

Selected expert outputs are aggregated with normalized weights:

$$\mathbf{y} = \sum_{i \in \mathcal{S}} w_i \cdot \mathrm{Expert}_{i,j_i^*}(\mathbf{x})$$

The auxiliary load-balancing loss takes the standard form:

$$\mathcal{L}_{\mathrm{aux}} = \alpha \cdot N \cdot \sum_{i=1}^{N} f_i \cdot p_i$$

where $f_i$ is the fraction of tokens routed to expert $i$, $p_i$ its mean routing probability, and $\alpha = 0.01$.

### 3.3 Training Pipeline

The pipeline has three stages, all using only public corpora:

1. **Stage 1 (domain-specialist initialization)**: Continually pre-trains four copies of the dense base on per-domain partitions of Dolmino-Mix (math, code, academic, web) for 4×10B = 40B tokens
2. **Stage 2 (joint continual pre-training)**: Trains the full DIVMOE model on Dolmino-Mix for 500B tokens; corpus, optimizer, and step count are identical for every baseline
3. **Stage 3 (supervised fine-tuning)**: Applies Tülu-3-SFT-Mixture identically to DIVMOE, NVIDIA-Upcycling, and OLMoE, isolating architectural effects from data effects

## Empirical Validation / Results

### Experimental Setup

- **Base models**: Qwen3-1.7B-Base and Llama3.2-1B (plus Qwen3-4B for the 12B MoE comparison)
- **Architecture**: $n=4$ domain groups, $m=2$ experts per group ($N=8$ total), $k=2$ active experts per token
- **Baselines**: From Scratch (FS), Sparse Upcycling (SU), Branch-Train-Mix (BTX), NVIDIA Upcycling (NU), Drop-Upcycling (DU), and its fine-grained variant (DU-f)
- **Evaluation**: 15 benchmarks spanning general knowledge, reasoning, code, and graduate-level science (ARC-c/e, BoolQ, COPA, MMLU, OpenBookQA, TriviaQA, HellaSwag, SQuAD2, GSM8K, MATH, BBH, HumanEval, MBPP, GPQA)

### Main Results on Qwen3-1.7B-Base (after 500B Stage 2 CPT tokens)

| Method | F.G. | ARC-c | BoolQ | MMLU | TriviaQ | HellaS. | GSM8K | MATH | HumanE. | MBPP | GPQA | BBH | Avg. |
|--------|------|-------|-------|------|---------|---------|-------|------|---------|------|------|-----|------|
| From Scratch | ✓ | 12.2 | 30.5 | 40.5 | 33.0 | 42.5 | 5.5 | 2.3 | 6.8 | 11.2 | 22.6 | 20.5 | 20.7 |
| Sparse Upcycling | ✗ | 38.5 | 60.2 | 56.8 | 43.5 | 54.5 | 56.5 | 28.5 | 38.0 | 47.5 | 24.5 | 47.5 | 45.1 |
| BTX | ✗ | 43.0 | 71.0 | 61.5 | 47.5 | 61.0 | 64.0 | 36.5 | 44.0 | 53.0 | 26.0 | 51.5 | 50.8 |
| NVIDIA Upcycling | ✓ | 42.5 | 70.5 | 61.5 | 47.5 | 60.5 | 62.5 | 35.5 | 44.5 | 53.5 | 25.8 | 51.0 | 50.5 |
| Drop-Upcycling | ✗ | 41.5 | 71.5 | 61.0 | 47.0 | 59.5 | 60.0 | 33.5 | 41.0 | 50.5 | 25.0 | 50.0 | 49.1 |
| Drop-Upcycling-f | ✓ | 14.5 | 32.5 | 42.0 | 33.5 | 41.5 | 7.0 | 3.5 | 7.8 | 12.5 | 22.5 | 22.0 | 21.8 |
| **DivMoE** | ✓ | **45.0** | **75.0** | **64.5** | **50.0** | **63.5** | **68.5** | **40.5** | **49.0** | **57.5** | **28.0** | **55.5** | **54.3** |

### Main Results on Llama3.2-1B (after 500B Stage 2 CPT tokens)

| Method | F.G. | ARC-c | BoolQ | MMLU | TriviaQ | HellaS. | GSM8K | MATH | HumanE. | MBPP | GPQA | BBH | Avg. |
|--------|------|-------|-------|------|---------|---------|-------|------|---------|------|------|-----|------|
| From Scratch | ✓ | 13.0 | 27.5 | 38.0 | 31.0 | 36.0 | 4.0 | 1.5 | 4.5 | 8.0 | 21.8 | 18.0 | 18.5 |
| Sparse Upcycling | ✗ | 33.5 | 60.5 | 48.0 | 39.5 | 52.5 | 26.5 | 9.0 | 15.0 | 21.5 | 23.5 | 31.5 | 32.8 |
| BTX | ✗ | 36.0 | 66.0 | 50.5 | 39.5 | 56.0 | 31.5 | 11.0 | 17.5 | 24.5 | 25.0 | 33.0 | 35.5 |
| NVIDIA Upcycling | ✓ | 36.0 | 65.5 | 50.0 | 39.0 | 55.5 | 30.5 | 10.8 | 17.0 | 24.0 | 24.8 | 32.8 | 35.1 |
| Drop-Upcycling | ✗ | 35.0 | 66.0 | 49.5 | 39.0 | 55.0 | 29.5 | 10.5 | 16.5 | 23.5 | 24.5 | 32.5 | 34.7 |
| Drop-Upcycling-f | ✓ | 25.5 | 33.0 | 43.5 | 38.0 | 39.5 | 6.0 | 2.5 | 6.0 | 9.5 | 21.5 | 19.0 | 22.2 |
| **DivMoE** | ✓ | **37.5** | **67.5** | **52.5** | **42.0** | **56.5** | **34.5** | **13.5** | **20.0** | **27.0** | **26.0** | **34.5** | **37.4** |

### Comparison with Open-Source MoE Models (Controlled SFT)

| Model | Size | ARC-c | MMLU | TriviaQA | HumanE. | MBPP | GSM8K | MATH | Avg. |
|-------|------|-------|------|----------|---------|------|-------|------|------|
| OLMoE | 1B/7B | 49.2 | 51.9 | 60.4 | 51.8 | 61.2 | 45.5 | 23.9 | 49.1 |
| OLMoE + Tülu-3 SFT | 1B/7B | 52.4 | 55.8 | 63.5 | 53.5 | 62.8 | 56.8 | 31.2 | 53.7 |
| DeepSeek-V2-Lite | 2.4B/16B | 52.1 | 58.3 | 65.1 | 29.9 | 43.2 | 41.1 | 17.1 | 43.8 |
| NVIDIA Up. + Tülu-3 SFT | 2.4B/5.4B | 56.3 | 64.2 | 67.5 | 45.8 | 54.2 | 66.5 | 38.8 | 56.2 |
| Moonlight-MoE | 2.4B/16B | 65.5 | 70.0 | 66.3 | 48.1 | 63.8 | 77.4 | 45.3 | 62.3 |
| **DivMoE** | 3.6B/12B | **63.1** | **71.5** | **73.3** | **51.4** | **60.0** | **74.4** | **47.5** | **63.0** |

### Ablation Studies

**Component synergy ablation** (Qwen3-1.7B-Base, 15 benchmarks):

| Configuration | Domain experts | Fine-grained | Diversity constraint | Avg. (15 bm.) |
|---------------|:--------------:|:------------:|:--------------------:|:-------------:|
| Drop-Upcycling-f | ✗ | ✓ | ✗ | 23.2 |
| NVIDIA Upcycling | ✗ | ✓ | ✗ | 51.1 |
| BTX | ✓ | ✗ | ✗ | 51.6 |
| DivMoE (- DC) | ✓ | ✓ | ✗ | 52.5 |
| **DivMoE** | ✓ | ✓ | ✓ | **55.6** |

**Effect of diversity constraint on domain-specific benchmarks**:

| Method | HumanE. | MBPP | GSM8K | MATH |
|--------|---------|------|-------|------|
| DIVMOE w/o DC | 23.8 | 31.5 | 43.2 | 18.1 |
| DIVMOE w/ DC | 25.6 | 33.8 | 45.3 | 19.5 |
| **Δ** | **+1.8** | **+2.3** | **+2.1** | **+1.4** |

### Key Findings

1. **Fine-grained collapse without domain diversity**: From-Scratch (20.7%) and Drop-Upcycling-fine-grained (21.8%) collapse to near-random performance on Qwen3-1.7B-Base, isolating the failure to the combination of fine-grained partitioning with non-domain-specialized initialization.

2. **Stage 2 CPT improves over the dense base**: DIVMOE surpasses Qwen3-1.7B-Base on every benchmark evaluated (e.g., MMLU 64.5 vs. 62.6, MATH 40.5 vs. 38.4, BBH 55.5 vs. 53.5), eliminating the regression gap that plagued prior upcycling pipelines.

3. **Routing patterns**: Sparse Upcycling shows routing collapse in deeper layers (over 60% of tokens routed to two experts at layer 36). With the constraint, routing is balanced across groups while preserving dataset-specific preferences: GSM8K activates more math experts, while HumanEval activates more code experts.

4. **Expert granularity**: Medium granularity ($m = k = 2$) provides the best balance between training-loss reduction and downstream accuracy; $m = k = 4$ shows mild overfitting on MMLU/HellaSwag.

## Theoretical and Practical Implications

### Theoretical Implications

1. **Routing collapse is structural, not just optimization-based**: The paper demonstrates that fine-grained upcycling collapse is a structural pathology, not merely a consequence of insufficient training or poor optimization. The collapse occurs regardless of the partitioning scheme when experts come from a single source, suggesting fundamental limitations in the representational diversity of single-source fine-grained MoEs.

2. **Hard constraints vs. soft losses**: The paper provides compelling evidence that hard structural constraints on routing (operating outside the loss landscape) can be more effective than soft auxiliary losses that trade routing balance against task accuracy. This challenges the prevailing paradigm of gradient-based load balancing in MoE training.

3. **Cross-domain composition is beneficial even for domain-specific tasks**: The counter-intuitive finding that forcing coding tokens to also draw from math, science, and commonsense experts *improves* coding performance suggests that individual tokens within a prompt serve heterogeneous functional roles that benefit from complementary expertise—a nuance lost in coarse-grained analyses of "domain-specific" tasks.

### Practical Implications

1. **Efficient MoE construction**: DIVMOE provides a practical recipe for converting existing dense models into fine-grained MoEs without the prohibitive cost of from-scratch training, achieving state-of-the-art performance with only 540B Stage-2 CPT tokens.

2. **Data-efficient scaling**: The 12B-parameter DIVMOE model matches Moonlight-MoE (16B, trained from scratch with substantially more compute), demonstrating that upcycling pipelines can rival much more expensive training regimes.

3. **Public corpora suffice**: All stages use only public corpora (Dolmino-Mix, Tülu-3-SFT-Mixture), making the entire pipeline reproducible and accessible to the research community.

4. **No regression after CPT**: The use of high-quality CPT corpora (Dolmino-Mix) eliminates the post-CPT regression that plagued prior upcycling pipelines, providing a template for future multi-stage training pipelines.

## Conclusion

DIVMOE is the first framework for fine-grained MoE upcycling that addresses the previously-unobserved fine-grained upcycling collapse pathology. The paper demonstrates that combining fine-grained expert partitioning with single-source initialization causes downstream accuracy to drop to near-random, irrespective of the partitioning scheme.

The framework resolves this issue through three complementary components:
1. **Domain-specialized fine-grained expert initialization**—deriving experts from $n$ dense models that have undergone domain-adaptive continual pre-training, then vertically slicing each specialist's FFN into $m$ shards
2. **Diversity-constrained routing**—a hard structural constraint guaranteeing that each token activates experts from distinct domain groups
3. **Public-corpus continual pre-training pipeline**—using Dolmino-Mix to eliminate post-CPT regression below the dense base

The ablation studies demonstrate that the proposed components are jointly necessary for robust fine-grained upcycling. Notably, the structural constraint improves performance on every domain-specific benchmark tested, refuting the intuition that cross-domain co-activation should harm specialized tasks.

**Future directions** suggested by this work include exploring larger numbers of domain groups, investigating the interaction between the diversity constraint and other routing mechanisms (e.g., expert-choice routing), and applying the framework to even larger base models and more diverse domain specializations.

---

_Markdown view of https://picx.dev/p/ZsznsA, served by PicX — AI-generated visual whiteboard summaries of research papers._
