DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
Summary (Overview)
-
Key Problem Identified: The paper uncovers a fine-grained upcycling collapse pathology—when fine-grained Mixture-of-Experts (MoE) experts are derived from a single source dense model, naive routing collapses and downstream accuracy drops to near-random levels (e.g., 23.2% average accuracy on Qwen3-1.7B, essentially matching from-scratch training at 22.2%).
-
Proposed Framework: DIVMOE is the first framework achieving fine-grained MoE Upcycling with structurally-balanced routing, combining two innovations: (1) domain-specialized fine-grained expert initialization from multiple domain-adapted dense models, and (2) diversity-constrained routing that structurally guarantees each token activates experts from distinct domain groups.
-
Performance Gains: DIVMOE consistently outperforms six upcycling baselines across two base models and 15 benchmarks (55.6% vs. 51.6% for the strongest baseline on Qwen3-1.7B), and strictly improves over the dense base model on every benchmark after Stage 2 continual pre-training.
-
Scalability Demonstration: After supervised fine-tuning, a 12B-parameter DIVMOE model matches Moonlight-MoE (16B) at 64.5% average accuracy while outperforming a controlled NVIDIA-Upcycling baseline by 6.1 percentage points.
-
Counter-Intuitive Finding: The structural diversity constraint improves performance on every domain-specific benchmark (HumanEval +1.8%, GSM8K +2.1%, MATH +1.4%, MBPP +2.3%), refuting the intuition that "coding tokens should only use coding experts."
Introduction and Theoretical Foundation
Background and Motivation
Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, enabling substantial capacity expansion while maintaining computational efficiency through sparse activation. Recent advances demonstrate the effectiveness of fine-grained MoE designs, where each feed-forward network (FFN) layer is partitioned into many small experts rather than a few large ones, contributing to the strong performance of models such as DeepSeek-V2 and Mixtral.
Training such models from scratch demands substantial compute, motivating sparse upcycling—converting pre-trained dense checkpoints into MoE architectures while inheriting their learned representations. However, the authors identify that fine-grained upcycling is fundamentally fragile when experts are derived from a single dense source.
The Collapse Pathology
The paper's central observation is that no prior method simultaneously occupies three design axes that are jointly necessary:
- (a) Domain-specialized expert initialization—experts derived from domain-adapted dense models
- (b) Fine-grained partitioning—FFN layers split into many small experts
- (c) A structural mechanism preventing same-source co-activation—hard constraints on routing
Existing methods fail on at least one axis:
- Branch Train-Mix (BTX) addresses (a) but uses coarse experts and standard top-k routing
- NVIDIA Upcycling addresses (b) but yields homogeneous expert initial scores that suppress routing diversity
- Drop-Upcycling targets diversity through randomness, sacrificing (a)
Theoretical Foundation
Given a dense transformer with SwiGLU FFN:
where and , the model converts each FFN into an MoE layer. With domain specialists, shards per specialist, and active experts, DIVMOE doubles total capacity while activating half the dense FFN parameters.
Methodology
3.1 Domain-Specialized Fine-Grained Expert Initialization
Domain specialists: Starting from base model , the method applies domain-adaptive continual pre-training (next-token-prediction loss on full sequences) on public domain-specific subsets to obtain specialists: . Each specialist encodes distinct domain knowledge while retaining general capabilities.
Expert construction: Each specialist's FFN weights are vertically sliced into shards:
where expert has dimensions for gate/up and for down projections. Indices and identify the specialist (domain group) and shard, respectively.
Dense-equivalent initialization: Following He et al., the method ensures stable initialization via virtual groups. Decomposing each specialist as where captures domain adaptations:
This preserves the base model's output distribution at initialization while introducing structured diversity through the terms.
Magnitude correction: Expert weights are scaled by to compensate for softmax attenuation of router scores:
where is the expansion factor, is granularity, and is the active-expert count.
Non-expert components: Embeddings, attention, and normalization layers are initialized as the mean of the base model and all specialists:
3.2 Diversity-Constrained Routing
Standard MoE routing computes scores and selects . When experts are partitioned into domain groups , this top-k rule admits within-group concentration.
The diversity constraint imposes that each token activates at most one expert per group:
combined with , guaranteeing that the selected experts originate from distinct domain groups.
Algorithm 1: Diversity-Constrained Routing
Input: hidden state x; router weights W_r; #groups n; experts/group m; active k
Output: selected experts S, weights w
s ← W_r x ▷ Router scores ∈ R^{n×m}
// Stage 1: Intra-group selection
for i = 1 to n do
j_i* ← argmax_{j ∈ [m]} s_{i,j} ▷ Best expert in group i
s_i* ← s_{i,j_i*} ▷ Representative score
end for
// Stage 2: Inter-group selection
G ← TopK({s_1*, ..., s_n*}, k) ▷ Top-k groups
S ← {(i, j_i*) : i ∈ G} ▷ Selected experts
// Compute normalized routing weights
for i ∈ G do
w_i ← exp(s_i*) / Σ_{j∈G} exp(s_j*)
end for
Return S, w
Selected expert outputs are aggregated with normalized weights:
The auxiliary load-balancing loss takes the standard form:
where is the fraction of tokens routed to expert , its mean routing probability, and .
3.3 Training Pipeline
The pipeline has three stages, all using only public corpora:
- Stage 1 (domain-specialist initialization): Continually pre-trains four copies of the dense base on per-domain partitions of Dolmino-Mix (math, code, academic, web) for 4×10B = 40B tokens
- Stage 2 (joint continual pre-training): Trains the full DIVMOE model on Dolmino-Mix for 500B tokens; corpus, optimizer, and step count are identical for every baseline
- Stage 3 (supervised fine-tuning): Applies Tülu-3-SFT-Mixture identically to DIVMOE, NVIDIA-Upcycling, and OLMoE, isolating architectural effects from data effects
Empirical Validation / Results
Experimental Setup
- Base models: Qwen3-1.7B-Base and Llama3.2-1B (plus Qwen3-4B for the 12B MoE comparison)
- Architecture: domain groups, experts per group ( total), active experts per token
- Baselines: From Scratch (FS), Sparse Upcycling (SU), Branch-Train-Mix (BTX), NVIDIA Upcycling (NU), Drop-Upcycling (DU), and its fine-grained variant (DU-f)
- Evaluation: 15 benchmarks spanning general knowledge, reasoning, code, and graduate-level science (ARC-c/e, BoolQ, COPA, MMLU, OpenBookQA, TriviaQA, HellaSwag, SQuAD2, GSM8K, MATH, BBH, HumanEval, MBPP, GPQA)
Main Results on Qwen3-1.7B-Base (after 500B Stage 2 CPT tokens)
| Method | F.G. | ARC-c | BoolQ | MMLU | TriviaQ | HellaS. | GSM8K | MATH | HumanE. | MBPP | GPQA | BBH | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| From Scratch | ✓ | 12.2 | 30.5 | 40.5 | 33.0 | 42.5 | 5.5 | 2.3 | 6.8 | 11.2 | 22.6 | 20.5 | 20.7 |
| Sparse Upcycling | ✗ | 38.5 | 60.2 | 56.8 | 43.5 | 54.5 | 56.5 | 28.5 | 38.0 | 47.5 | 24.5 | 47.5 | 45.1 |
| BTX | ✗ | 43.0 | 71.0 | 61.5 | 47.5 | 61.0 | 64.0 | 36.5 | 44.0 | 53.0 | 26.0 | 51.5 | 50.8 |
| NVIDIA Upcycling | ✓ | 42.5 | 70.5 | 61.5 | 47.5 | 60.5 | 62.5 | 35.5 | 44.5 | 53.5 | 25.8 | 51.0 | 50.5 |
| Drop-Upcycling | ✗ | 41.5 | 71.5 | 61.0 | 47.0 | 59.5 | 60.0 | 33.5 | 41.0 | 50.5 | 25.0 | 50.0 | 49.1 |
| Drop-Upcycling-f | ✓ | 14.5 | 32.5 | 42.0 | 33.5 | 41.5 | 7.0 | 3.5 | 7.8 | 12.5 | 22.5 | 22.0 | 21.8 |
| DivMoE | ✓ | 45.0 | 75.0 | 64.5 | 50.0 | 63.5 | 68.5 | 40.5 | 49.0 | 57.5 | 28.0 | 55.5 | 54.3 |
Main Results on Llama3.2-1B (after 500B Stage 2 CPT tokens)
| Method | F.G. | ARC-c | BoolQ | MMLU | TriviaQ | HellaS. | GSM8K | MATH | HumanE. | MBPP | GPQA | BBH | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| From Scratch | ✓ | 13.0 | 27.5 | 38.0 | 31.0 | 36.0 | 4.0 | 1.5 | 4.5 | 8.0 | 21.8 | 18.0 | 18.5 |
| Sparse Upcycling | ✗ | 33.5 | 60.5 | 48.0 | 39.5 | 52.5 | 26.5 | 9.0 | 15.0 | 21.5 | 23.5 | 31.5 | 32.8 |
| BTX | ✗ | 36.0 | 66.0 | 50.5 | 39.5 | 56.0 | 31.5 | 11.0 | 17.5 | 24.5 | 25.0 | 33.0 | 35.5 |
| NVIDIA Upcycling | ✓ | 36.0 | 65.5 | 50.0 | 39.0 | 55.5 | 30.5 | 10.8 | 17.0 | 24.0 | 24.8 | 32.8 | 35.1 |
| Drop-Upcycling | ✗ | 35.0 | 66.0 | 49.5 | 39.0 | 55.0 | 29.5 | 10.5 | 16.5 | 23.5 | 24.5 | 32.5 | 34.7 |
| Drop-Upcycling-f | ✓ | 25.5 | 33.0 | 43.5 | 38.0 | 39.5 | 6.0 | 2.5 | 6.0 | 9.5 | 21.5 | 19.0 | 22.2 |
| DivMoE | ✓ | 37.5 | 67.5 | 52.5 | 42.0 | 56.5 | 34.5 | 13.5 | 20.0 | 27.0 | 26.0 | 34.5 | 37.4 |
Comparison with Open-Source MoE Models (Controlled SFT)
| Model | Size | ARC-c | MMLU | TriviaQA | HumanE. | MBPP | GSM8K | MATH | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| OLMoE | 1B/7B | 49.2 | 51.9 | 60.4 | 51.8 | 61.2 | 45.5 | 23.9 | 49.1 |
| OLMoE + Tülu-3 SFT | 1B/7B | 52.4 | 55.8 | 63.5 | 53.5 | 62.8 | 56.8 | 31.2 | 53.7 |
| DeepSeek-V2-Lite | 2.4B/16B | 52.1 | 58.3 | 65.1 | 29.9 | 43.2 | 41.1 | 17.1 | 43.8 |
| NVIDIA Up. + Tülu-3 SFT | 2.4B/5.4B | 56.3 | 64.2 | 67.5 | 45.8 | 54.2 | 66.5 | 38.8 | 56.2 |
| Moonlight-MoE | 2.4B/16B | 65.5 | 70.0 | 66.3 | 48.1 | 63.8 | 77.4 | 45.3 | 62.3 |
| DivMoE | 3.6B/12B | 63.1 | 71.5 | 73.3 | 51.4 | 60.0 | 74.4 | 47.5 | 63.0 |
Ablation Studies
Component synergy ablation (Qwen3-1.7B-Base, 15 benchmarks):
| Configuration | Domain experts | Fine-grained | Diversity constraint | Avg. (15 bm.) |
|---|---|---|---|---|
| Drop-Upcycling-f | ✗ | ✓ | ✗ | 23.2 |
| NVIDIA Upcycling | ✗ | ✓ | ✗ | 51.1 |
| BTX | ✓ | ✗ | ✗ | 51.6 |
| DivMoE (- DC) | ✓ | ✓ | ✗ | 52.5 |
| DivMoE | ✓ | ✓ | ✓ | 55.6 |
Effect of diversity constraint on domain-specific benchmarks:
| Method | HumanE. | MBPP | GSM8K | MATH |
|---|---|---|---|---|
| DIVMOE w/o DC | 23.8 | 31.5 | 43.2 | 18.1 |
| DIVMOE w/ DC | 25.6 | 33.8 | 45.3 | 19.5 |
| Δ | +1.8 | +2.3 | +2.1 | +1.4 |
Key Findings
-
Fine-grained collapse without domain diversity: From-Scratch (20.7%) and Drop-Upcycling-fine-grained (21.8%) collapse to near-random performance on Qwen3-1.7B-Base, isolating the failure to the combination of fine-grained partitioning with non-domain-specialized initialization.
-
Stage 2 CPT improves over the dense base: DIVMOE surpasses Qwen3-1.7B-Base on every benchmark evaluated (e.g., MMLU 64.5 vs. 62.6, MATH 40.5 vs. 38.4, BBH 55.5 vs. 53.5), eliminating the regression gap that plagued prior upcycling pipelines.
-
Routing patterns: Sparse Upcycling shows routing collapse in deeper layers (over 60% of tokens routed to two experts at layer 36). With the constraint, routing is balanced across groups while preserving dataset-specific preferences: GSM8K activates more math experts, while HumanEval activates more code experts.
-
Expert granularity: Medium granularity () provides the best balance between training-loss reduction and downstream accuracy; shows mild overfitting on MMLU/HellaSwag.
Theoretical and Practical Implications
Theoretical Implications
-
Routing collapse is structural, not just optimization-based: The paper demonstrates that fine-grained upcycling collapse is a structural pathology, not merely a consequence of insufficient training or poor optimization. The collapse occurs regardless of the partitioning scheme when experts come from a single source, suggesting fundamental limitations in the representational diversity of single-source fine-grained MoEs.
-
Hard constraints vs. soft losses: The paper provides compelling evidence that hard structural constraints on routing (operating outside the loss landscape) can be more effective than soft auxiliary losses that trade routing balance against task accuracy. This challenges the prevailing paradigm of gradient-based load balancing in MoE training.
-
Cross-domain composition is beneficial even for domain-specific tasks: The counter-intuitive finding that forcing coding tokens to also draw from math, science, and commonsense experts improves coding performance suggests that individual tokens within a prompt serve heterogeneous functional roles that benefit from complementary expertise—a nuance lost in coarse-grained analyses of "domain-specific" tasks.
Practical Implications
-
Efficient MoE construction: DIVMOE provides a practical recipe for converting existing dense models into fine-grained MoEs without the prohibitive cost of from-scratch training, achieving state-of-the-art performance with only 540B Stage-2 CPT tokens.
-
Data-efficient scaling: The 12B-parameter DIVMOE model matches Moonlight-MoE (16B, trained from scratch with substantially more compute), demonstrating that upcycling pipelines can rival much more expensive training regimes.
-
Public corpora suffice: All stages use only public corpora (Dolmino-Mix, Tülu-3-SFT-Mixture), making the entire pipeline reproducible and accessible to the research community.
-
No regression after CPT: The use of high-quality CPT corpora (Dolmino-Mix) eliminates the post-CPT regression that plagued prior upcycling pipelines, providing a template for future multi-stage training pipelines.
Conclusion
DIVMOE is the first framework for fine-grained MoE upcycling that addresses the previously-unobserved fine-grained upcycling collapse pathology. The paper demonstrates that combining fine-grained expert partitioning with single-source initialization causes downstream accuracy to drop to near-random, irrespective of the partitioning scheme.
The framework resolves this issue through three complementary components:
- Domain-specialized fine-grained expert initialization—deriving experts from dense models that have undergone domain-adaptive continual pre-training, then vertically slicing each specialist's FFN into shards
- Diversity-constrained routing—a hard structural constraint guaranteeing that each token activates experts from distinct domain groups
- Public-corpus continual pre-training pipeline—using Dolmino-Mix to eliminate post-CPT regression below the dense base
The ablation studies demonstrate that the proposed components are jointly necessary for robust fine-grained upcycling. Notably, the structural constraint improves performance on every domain-specific benchmark tested, refuting the intuition that cross-domain co-activation should harm specialized tasks.
Future directions suggested by this work include exploring larger numbers of domain groups, investigating the interaction between the diversity constraint and other routing mechanisms (e.g., expert-choice routing), and applying the framework to even larger base models and more diverse domain specializations.
Related papers
- Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents
Tool calls are silently altered by execution-path hops in 12% of production shell invocations, and the IntAct repair protocol recovers 79.2% of resulting failures.
- TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
TESTJACK refutes 34.4% of benchmark-passing LLM coding trials by evolving tests from prompt requirements, cutting reported resolution rates from 50.6% to 33.2%.
- More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding
SAGA decouples key and value head counts to exploit sparse attention's shifted bottleneck, achieving over 2x decoding speedup at 128K context with near-baseline quality.