Full text not available for this paper
Summary (Overview)
- This paper investigates how the optimal learning-rate warmup duration scales with the training horizon in language model training, showing that it is neither fixed nor a constant fraction of the horizon but follows a regime-dependent scaling law.
- The authors identify three distinct regimes: essentially no warmup at low peak learning rates, bounded warmup at intermediate rates, and warmup durations that grow (sometimes near-proportionally) with the training horizon at high peak rates.
- A simple quadratic model with two modes (low-curvature and near-stability) explains the tradeoff: warmup sacrifices early progress in well-conditioned directions but removes persistent error in directions near the stability edge, with higher peak rates shifting the balance toward longer warmup.
- The paper derives a compact loss law that captures the warmup-horizon tradeoff and can be fit using as few as three short training runs to predict optimal warmup at substantially longer horizons.
- The findings suggest treating warmup duration as a horizon-dependent hyperparameter that should be tuned jointly with peak learning rate and training budget, rather than as a fixed heuristic.
Introduction and Theoretical Foundation
Background and Motivation
Learning-rate warmup is ubiquitous in language-model training, yet its duration remains largely heuristic. Common approaches use either:
- Fixed number of warmup steps (Vaswani et al., 2017; Groeneveld et al., 2024; Kaplan et al., 2020)
- Fixed fraction of the training horizon (Black et al., 2022)
These choices imply very different scaling as training gets longer, raising the question: When should warmup stay fixed, and when should it grow with the horizon?
Prior Work
Previous explanations of warmup include:
- Adaptive-optimizer statistics and update magnitudes (Liu et al., 2020; Ma & Yarats, 2021)
- Loss curvature and access to larger subsequent learning rates (Gilmer et al., 2022; Kalra & Barkeshli, 2024)
- Early changes in parameters and representations (Gotmare et al., 2019; Kosson et al., 2024)
- Suboptimality-dependent smoothness (Liu et al., 2025b; Alimisis et al., 2026; Riabinin et al., 2026)
However, explaining why warmup helps does not determine its optimal duration. That requires balancing the persistent benefit of longer warmup against the progress lost to smaller initial steps.
Theoretical Foundation: Quadratic Model
The paper uses a quadratic model as a standard lens for studying optimization effects:
In a direction of curvature , one update with learning rate multiplies that direction's loss by . Defining , the contraction rate at the peak learning rate and its average over linear warmup are:
The loss remaining after warmup and the subsequent peak-rate phase is:
Methodology
Experimental Design
- Models: Llama-style language models of sizes 60M, 100M, and 350M parameters
- Training: Pretrained on C4 English with t5-base SentencePiece tokenizer, sequence length 256, batch size 512 (131,072 tokens per update)
- Schedule: Linear warmup followed by constant peak rate (warmup–stable), which allows one trajectory to be evaluated at many horizons
- Grid: 33 model-size and peak-rate families, with warmup ranges from 0 to 64,000 updates
- Peak learning rates: 0.1 to 8 × 10⁻³
The Loss Law
The paper derives a compact loss law over warmup duration and training horizon :
where:
- is the effective peak-rate-equivalent training amount (a warmup step counts as half a peak-rate step)
- is the irreducible loss
- The term captures the cost of delayed progress (longer warmup reduces )
- The term captures the persistent error reduction from warmup
- (in thousands of updates) keeps the predictor well-defined at zero warmup
Selection Rule
For warmup selection, only differences between candidate losses matter:
Empirical Validation / Results
Key Empirical Findings
1. Peak learning rate determines warmup regime:
| Regime | Peak LR | Behavior |
|---|---|---|
| Conservative | Low (e.g., 0.0003) | Little or no warmup; preferred duration changes little with horizon |
| Aggressive | High (e.g., 0.007) | Optimal warmup grows substantially with training horizon |
2. Ranking reversal: At high peak rates, a longer warmup can start worse but become better later, so its benefit persists beyond the warmup duration itself.
3. Best runs use long warmups (Table 1):
| Model | Peak LR | Warmup (k) | Fraction of horizon (T=128k) | Loss |
|---|---|---|---|---|
| 60M | 0.004 | 30 | 23.4% | 3.404 |
| 100M | 0.005 | 50 | 39.1% | 3.246 |
| 350M | 0.004 | 40 | 31.3% | 2.983 |
4. Shortest successful warmup increases with:
- Higher peak learning rates
- Smaller batch sizes
- Larger models
Theoretical Results
Two-mode optimum: With a low-curvature mode (where ) and a near-stability mode (where ), the optimal warmup is:
If and , the long-horizon optimum changes at :
- Zero for
- Bounded at
- A fraction of T for , growing with
Fit Quality and Prediction
Table 2: Absolute-fit accuracy and selection quality
| Evaluation | Median R² | RMSE | Mean regret |
|---|---|---|---|
| All horizons (descriptive) | 0.996 | 4.92 | 2.00 ± 1.51 |
| Alternating holdout | 0.996 | 4.98 | 2.05 ± 1.61 |
RMSE and regret in 10⁻³ loss units.
Warmup Growth Exponent
The fitted growth exponent from the loss law predicts:
- : zero or bounded warmup
- : sublinear growth
- : proportional growth
Extrapolation Results
Table 3: Fits through 32k select low-regret warmups at later horizons
| Selector | Parameters | Alternating | Four horizons | 128k |
|---|---|---|---|---|
| Absolute fit | 6 | 2.05 ± 1.61 | 2.26 ± 0.66 | 3.05 |
| Difference fit | 5 | 1.45 ± 2.22 | 1.58 ± 0.77 | 2.61 |
| 1k target | 0 | 15.47 ± 10.80 | 13.79 ± 0.22 | 13.56 |
| 10% target | 0 | 13.39 ± 4.10 | 12.58 ± 0.90 | 11.72 |
Mean regret in 10⁻³ loss units.
Table 4: Three short runs support longer-horizon warmup prediction
| Fit horizon | Absolute | Difference | Fixed dur. | Fixed frac. | 1k | 10% |
|---|---|---|---|---|---|---|
| 20k (29 families) | 2.83 ± 0.16 | 4.22 ± 0.29 | 9.31 ± 1.21 | 11.72 ± 0.82 | 14.28 ± 0.30 | 13.42 ± 0.96 |
| 32k (33 families) | 2.51 ± 0.27 | 2.62 ± 0.36 | 7.61 ± 1.12 | 8.47 ± 0.62 | 13.79 ± 0.22 | 12.58 ± 0.90 |
Mean regret over 50k, 75k, 100k, and 128k, in 10⁻³ loss units.
Long-Horizon Extrapolation (up to 1M updates)
Table 5: Long-horizon validation-loss comparison (100M model, batch size 128)
| LR = 0.5×10⁻³ | LR = 2×10⁻³ | LR = 5×10⁻³ | |||||||
|---|---|---|---|---|---|---|---|---|---|
| T (k) | 250 | 500 | 1000 | 250 | 500 | 1000 | 250 | 500 | 1000 |
| 1k-warmup loss | 3.370 | 3.337 | 3.311 | 3.433 | 3.421 | 3.402 | 3.371 | 3.347 | 3.322 |
| 10% loss | 3.375 | 3.347 | 3.339 | 3.418 | 3.367 | 3.340 | 3.363 | 3.335 | 3.314 |
| Ŵ₅₀ | 0.6 | 0.7 | 0.8 | 47.7 | 77.5 | 125.9 | 129.5 | 241.7 | 451.1 |
| Loss | 3.372 | 3.336 | 3.311 | 3.412 | 3.362 | 3.338 | 3.359 | 3.325 | 3.303 |
| Ŵ₇₅ | 0.8 | 1.1 | 1.2 | 56.6 | 95.2 | 160.0 | 116.0 | 213.5 | 392.9 |
| Loss | 3.371 | 3.336 | 3.312 | 3.411 | 3.359 | 3.336 | 3.358 | 3.324 | 3.303 |
Horizons and predicted warmups in thousands of updates.
Theoretical and Practical Implications
Theoretical Implications
-
Unified explanation: The quadratic model explains several familiar properties of warmup through a single tradeoff—between giving up early progress and reducing persistent error near the stability edge.
-
Stability requirements: The model predicts the minimum warmup fraction to avoid divergence:
which increases with the peak rate and eventually reaches one.
- Scaling law framework: Warmup duration joins other optimization hyperparameters (learning rate, batch size, weight decay) that have their own scaling laws with training budget.
Practical Implications
-
Warmup should be tuned jointly with peak learning rate and training budget—fixing warmup in advance can change which learning rate appears best.
-
Practical prediction protocol: Fit the loss law using only three short runs (warmups of 2k, 8k, and 16k updates) with checkpoints through 20k or 32k updates, then predict warmup at substantially longer horizons with low regret.
-
Regime awareness: Practitioners should expect:
- Low peak rates → short or bounded warmup
- High peak rates → warmup that grows with the horizon
-
Robustness: The warmup ordering is largely preserved when decay is appended (warmup-stable-decay schedules), and the behavior holds across different warmup shapes (linear, half-cosine, concave-quadratic) and optimizers (AdamW, Muon).
Conclusion
The paper demonstrates that warmup duration is neither fixed across training budgets nor tied to a fixed fraction of the run. The key findings are:
-
Regime-dependent scaling: Lower peak learning rates favor short or bounded warmup, while higher rates favor durations that grow with the training horizon.
-
Mechanistic explanation: A simple quadratic model explains this shift as a balance between sacrificing early progress and reducing error that would otherwise persist later in training.
-
Compact loss law: The derived loss law captures the observed warmup regimes and can be fit using only three short trajectories to guide warmup choices at longer horizons.
-
Practical recommendation: Treat warmup duration as a horizon-dependent hyperparameter, tuned jointly with peak learning rate and training budget, rather than as a fixed training heuristic.
Future Directions
- Model-scale extrapolation: The authors show a correspondence between model size and effective peak learning rate, suggesting a model-size correction could extend predictions across scales.
- Understanding how warmup interacts with other schedule components (decay phase, annealing) in more complex schedules.
- Extending the analysis to other optimizers and training regimes beyond the ones studied.
Related papers
- How Local Mixing Encodes Relative Position in Global NoPE Attention
Hybrid architectures with local mixing layers and NoPE global attention implicitly learn relative position encodings via recency bias, enabling superior length extrapolation.
- Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
SuffixReplay enables fine-grained prefix caching in hybrid LLMs by reconstructing linear-attention states from sparse anchors, cutting storage by 2x and TTFT by up to 70%.
- Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers
Newton–Schulz iterations expose spectral moments via free scalar reductions, enabling spectrum-adaptive routine selection that cuts polar error up to 90x and improves LLM pretraining loss.