Warmup Duration Enters Scaling Laws: Hyperparameters Shift from Fixed Heuristics to Predictable Axes
Highlights of This Issue
The most prominent progress this issue is the re-integration of hyperparameters previously treated as fixed heuristics into the scaling law framework. Balancing Early Performance Sacrifices with Long-Term Gains (https://arxiv.org/abs/2609.33041) is the first to treat the learning rate warmup duration W itself as a hyperparameter that scales with training level T, proposing a compact loss law L(W,T)=L∞+Aτ^{-p}+Kτ^{-q}(W+w0)^{-s}, and proving that the optimal warmup duration grows as a power law with T, with the growth exponent determined by the peak learning rate. This extends the discussion on Step Law calibration and extrapolation reliability from Issue 7 to the warmup axis, previously treated as a fixed detail, and provides a practical protocol for extrapolating longer levels from three short runs. Complementing this, How Does Local Landscape Geometry Evolve (https://arxiv.org/abs/2609.39767) offers a geometric explanation for "large peak LR requires longer warmup" from the two-stage evolution of the loss landscape, with the two lines mutually corroborating at a mechanistic level.
On the optimizer side, spectral information shifts from offline measurement to online adaptation. Cost-free Spectral Estimation (https://arxiv.org/abs/2609.33047) is the first to extract spectral moments at zero cost from the intermediate Gram matrices of Newton–Schulz iterations, using maximum entropy moment fitting to recover the empirical singular value distribution and adaptively selecting polar decomposition routines for each matrix, directly addressing the open question of "how to cheaply obtain spectral information" left by Spectral Scaling Laws of Muon in Issue 3 and Spectral Allocation in Issue 4. Normalize-Then-Precondition (https://arxiv.org/abs/2609.36692) decomposes Muon's joint geometric transformation into a two-level hierarchy of "marginal norm normalization + spectral preconditioning," proposing global (NormPre-G) and local (NormPre-L) spectral preconditioning implementations, providing convergence guarantees and efficiency-performance trade-offs at 124M–1.7B. SOLAR (https://arxiv.org/abs/2609.34681) is the first to successfully apply an online learning rate scheduler (L2O/RL) to the full LLM pretraining process, ensuring stability through residual modulation and circuit breakers, and supporting frozen transfer across scales.
Pretraining objectives and scaling law methodology also see progress. On Trajectory-Aware Training (https://arxiv.org/abs/2609.37974) unifies trajectory-aware training for masked diffusion (progressive unmasking + carry + BPTT) into the PUMBA framework, revealing local overfitting under small u, and matching or surpassing longer SFT with fewer NFE on LLaDA-8B SFT, directly addressing the data efficiency debate between masked diffusion and autoregressive models. ScAn-Bench (https://arxiv.org/abs/2609.35707) is the first to turn scaling law methodology itself into an interactive surrogate benchmark, fitting TabPFN surrogates with 4524/8024 checkpoints, systematically revealing that CARBS's local exploitation leads to N* extrapolation failure and that Chinchilla parameterization collapses on VLMs due to the failure of the C≈6ND assumption. How Far Are We from Removing the Visual Encoder (https://arxiv.org/abs/2609.35457) is the first to systematically apply scaling law methods to compare encoder-free and encoder-based multimodal architectures, finding that removing the visual encoder shifts the compute-optimal allocation of multimodal objectives toward larger models.
Community and Updates
- The Death of Chinchilla (https://rickxie.cn/blog/pretraining-the-death-of-chinchilla/) challenges the applicability of the 2026 Chinchilla compute-optimal allocation in a position paper, framing the gap between "smooth loss extrapolation and unpredictable capability emergence" as an open problem, serving as a timely debate around compute-optimal scaling laws.
- Diffusion vs autoregressive, at nanoGPT scale (https://kareemfareed.com/posts/diffusion-vs-ar/) provides a compute-matched comparison of masked diffusion and autoregressive models at the nanoGPT scale, clearly separating compute-matched and data-matched conditions, directly countering claims about masked diffusion's data efficiency crossover.
- ICML 2026's How to Scale Mixture-of-Experts (https://icml.cc/virtual/2026/72343) formally incorporates the muP-for-MoE fix from Issue 1's MSSP, diagnosing the failure of muP assumptions under sparse expert routing and providing scale-stable parameterization.
- Let's Scale Step by Step (https://sparsenotes.com/posts/2026/08/papers/mup-hyperparameter-transfer-moe/) community interpretation explicitly notes that muP width transfer does not cover the token-horizon axis, identifying the intersection of token budget scaling laws and width transfer as an open problem for sparse MoE.
Open Questions
- The optimal warmup duration exponent β=(p+1-q)/(s+1) in the warmup scaling law is determined by the peak learning rate, but Issue 7's Step Law calibration has shown that power law coefficients shift significantly at small scales. Does the extrapolation reliability of the warmup exponent also drift with scale, or is it more robust than the learning rate itself?
- Can the online spectral moment extraction from Cost-free Spectral Estimation and the per-rank optimal step size spectrum from Issue 4's Spectral Allocation be unified—does online adaptive selection converge to the same "head-anchored, body-amplified" spectral prior?
- Is the local overfitting revealed by PUMBA specific to data-limited SFT scenarios, or does it also exist in masked diffusion pretraining itself? Can trajectory-aware training bring compute-matched gains at pretraining scale, rather than only in the SFT phase?
- The surrogate benchmark ceiling of ScAn-Bench is limited to LLM 1B / VLM 357M; do the CARBS local exploitation failure and Chinchilla parameterization collapse it reveals hold at larger scales, or does surrogate fitting itself obscure real training dynamics?
Papers in this issue
Optimal learning-rate warmup duration scales with training horizon only at high peak learning rates, following a regime-dependent law predictable from three short runs.
Editor's noteFirst to elevate the learning rate warmup duration from a fixed heuristic to a hyperparameter that scales with training level, proposing a compact loss law involving W and T and proving that the optimal warmup duration grows as a power law with T, with the growth exponent determined by the peak learning rate. Validated at 60M/100M/350M scales and providing a protocol for extrapolation from three short runs, directly addressing the extrapolation reliability of hyperparameter scaling laws.
Newton–Schulz iterations expose spectral moments via free scalar reductions, enabling spectrum-adaptive routine selection that cuts polar error up to 90x and improves LLM pretraining loss.
Editor's noteFirst to extract spectral moments at zero cost from the intermediate Gram matrices of Newton–Schulz iterations, using maximum entropy moment fitting to recover the empirical singular value distribution and adaptively selecting polar decomposition routines for each matrix. Compared to the static spectral scaling laws of Spectral Scaling Laws of Muon in Issue 3 and the offline measurements of Spectral Allocation in Issue 4, it turns spectral information into an online, per-matrix, training-adaptive mechanism, validated at 160M–1B scales.
PUMBA trains masked diffusion language models on inference-like trajectories via continuous hidden-state carries and backpropagation through time, matching autoregressive accuracy while decoding multiple tokens per step.
Editor's noteUnifies trajectory-aware training for masked diffusion (progressive unmasking + carry + BPTT) into the PUMBA framework, systematically revealing local overfitting under small u, proving that continuous carry outperforms discrete gradient estimators and that longer BPTT windows are better. Matches or surpasses longer SFT with fewer NFE on LLaDA-8B SFT, directly addressing the data efficiency debate between masked diffusion and autoregressive models.
ScAn-Bench introduces the first surrogate benchmarks for scaling analysis, revealing no universally optimal data acquisition and extrapolation strategy exists across LLMs and VLMs.
Editor's noteFirst work to turn scaling law methodology itself into an interactive surrogate benchmark, fitting TabPFN surrogates with 4524/8024 checkpoints and systematically evaluating the interaction effects of data acquisition and extrapolation strategies. Finds that CARBS's local exploitation leads to N* extrapolation failure and that Chinchilla parameterization collapses on VLMs due to the failure of the C≈6ND assumption, providing reproducible comparative evidence for scaling law fitting.
SOLAR uses reinforcement learning to learn bounded residual corrections to a base learning-rate schedule, improving LLM pretraining perplexity across dense and MoE models up to 3B parameters.
Editor's noteFirst to successfully apply an online learning rate scheduler (L2O/RL) to the full LLM pretraining process, ensuring stability through residual modulation decoupled from the base scheduler and circuit breakers, validated at 60M–1B and 3B MoE scales, and supporting frozen transfer across scales. Compared to static/analytic schedulers, it elevates learning rate scheduling from heuristics to online adaptation.
Encoder-free multimodal LLMs match encoder-based performance at ~10^22 FLOPs, shifting compute-optimal allocation toward larger decoders and enabling viable encoder-free pretraining.
Editor's noteFirst to systematically apply scaling law methods to compare encoder-free and encoder-based multimodal architectures, finding that removing the visual encoder shifts the compute-optimal allocation of multimodal objectives toward larger models (a rises from 0.464 to 0.570), and predicting an efficiency crossover at approximately 10^22 FLOPs. Validated at 1.1B–44B MoE scales with reported extrapolation residuals and bootstrap uncertainties.
NormPre separates LLM update normalization from spectral preconditioning, consistently beating AdamW, Muon, and MANO in pretraining while cutting optimizer latency by up to 67%.
Editor's noteDecomposes Muon's joint geometric transformation into a two-level hierarchy of "marginal norm normalization + spectral preconditioning," proposing global (NormPre-G) and local (NormPre-L) spectral preconditioning implementations, with the local version using random sketching to only clip super-unit dominant modes. Provides convergence guarantees and consistent improvements over AdamW/Muon/MANO at 124M–1.7B, serving as a systematic framework for Muon-family optimizers.
Fast learning-rate transfer holds at growing training horizons when T grows slower than sqrt(n), but requires nondegenerate first-order loss sensitivity to avoid spectral-dependent failures.
Editor's noteStrictly characterizes the rate and limit distribution of learning rate transfer when training duration T and width n grow jointly under shallow linear networks, μP initialization, and full-batch gradient descent, providing a sufficient condition for T=o(√n) and revealing that the transfer rate is jointly determined by three time-varying scales: loss response, gradient response, and local curvature. Provides rigorous theory for the reliability of muP transfer as training duration grows.







