Tracked direction
MoE & Sparse Experts
This direction focuses on routing mechanisms and system implementations in sparse mixture-of-experts (MoE) models, covering train-inference routing consistency, load-balancing algorithms, empirical tests of expert specialization, scaling behavior of sparsity and expert granularity, and expert parallelism with serving deployment. The emphasis is on validation at frontier scales, quantitative reporting of routing statistics and load distributions, and comparisons against dense baselines, aiming to understand and optimize the behavior and efficiency of sparse expert structures themselves.