Full text not available for this paper

Summary (Overview)

  • This paper investigates how the optimal learning-rate warmup duration scales with the training horizon in language model training, showing that it is neither fixed nor a constant fraction of the horizon but follows a regime-dependent scaling law.
  • The authors identify three distinct regimes: essentially no warmup at low peak learning rates, bounded warmup at intermediate rates, and warmup durations that grow (sometimes near-proportionally) with the training horizon at high peak rates.
  • A simple quadratic model with two modes (low-curvature and near-stability) explains the tradeoff: warmup sacrifices early progress in well-conditioned directions but removes persistent error in directions near the stability edge, with higher peak rates shifting the balance toward longer warmup.
  • The paper derives a compact loss law L(W,T)=L∞+Aτ−p+Kτ−q(W+w0)−sL(W,T) = L_\infty + A\tau^{-p} + K\tau^{-q}(W + w_0)^{-s} that captures the warmup-horizon tradeoff and can be fit using as few as three short training runs to predict optimal warmup at substantially longer horizons.
  • The findings suggest treating warmup duration as a horizon-dependent hyperparameter that should be tuned jointly with peak learning rate and training budget, rather than as a fixed heuristic.

Introduction and Theoretical Foundation

Background and Motivation

Learning-rate warmup is ubiquitous in language-model training, yet its duration remains largely heuristic. Common approaches use either:

  • Fixed number of warmup steps (Vaswani et al., 2017; Groeneveld et al., 2024; Kaplan et al., 2020)
  • Fixed fraction of the training horizon (Black et al., 2022)

These choices imply very different scaling as training gets longer, raising the question: When should warmup stay fixed, and when should it grow with the horizon?

Prior Work

Previous explanations of warmup include:

  • Adaptive-optimizer statistics and update magnitudes (Liu et al., 2020; Ma & Yarats, 2021)
  • Loss curvature and access to larger subsequent learning rates (Gilmer et al., 2022; Kalra & Barkeshli, 2024)
  • Early changes in parameters and representations (Gotmare et al., 2019; Kosson et al., 2024)
  • Suboptimality-dependent smoothness (Liu et al., 2025b; Alimisis et al., 2026; Riabinin et al., 2026)

However, explaining why warmup helps does not determine its optimal duration. That requires balancing the persistent benefit of longer warmup against the progress lost to smaller initial steps.

Theoretical Foundation: Quadratic Model

The paper uses a quadratic model as a standard lens for studying optimization effects:

f(θ)−f⋆=12(θ−θ⋆)⊤H(θ−θ⋆),H⪰0f(\theta) - f^\star = \frac{1}{2}(\theta - \theta^\star)^\top H(\theta - \theta^\star), \quad H \succeq 0

In a direction of curvature h>0h > 0, one update with learning rate ηt\eta_t multiplies that direction's loss by ∣1−ηth∣2|1 - \eta_t h|^2. Defining α=ηh\alpha = \eta h, the contraction rate at the peak learning rate and its average over linear warmup are:

ρ(α)=−2log⁡∣1−α∣,g(α)=∫01ρ(αu) du\rho(\alpha) = -2\log|1-\alpha|, \quad g(\alpha) = \int_0^1 \rho(\alpha u)\,du

The loss remaining after warmup and the subsequent peak-rate phase is:

MT,W(α)≈e−g(α)W−ρ(α)(T−W)M_{T,W}(\alpha) \approx e^{-g(\alpha)W - \rho(\alpha)(T-W)}

Methodology

Experimental Design

  • Models: Llama-style language models of sizes 60M, 100M, and 350M parameters
  • Training: Pretrained on C4 English with t5-base SentencePiece tokenizer, sequence length 256, batch size 512 (131,072 tokens per update)
  • Schedule: Linear warmup followed by constant peak rate (warmup–stable), which allows one trajectory to be evaluated at many horizons
  • Grid: 33 model-size and peak-rate families, with warmup ranges from 0 to 64,000 updates
  • Peak learning rates: 0.1 to 8 × 10⁻³

The Loss Law

The paper derives a compact loss law over warmup duration WW and training horizon TT:

L(W,T)=L∞+Aτ−p+Kτ−q(W+w0)−sL(W, T) = L_\infty + A\tau^{-p} + K\tau^{-q}(W + w_0)^{-s}

where:

  • τ=T−W/2\tau = T - W/2 is the effective peak-rate-equivalent training amount (a warmup step counts as half a peak-rate step)
  • L∞L_\infty is the irreducible loss
  • The Aτ−pA\tau^{-p} term captures the cost of delayed progress (longer warmup reduces τ\tau)
  • The Kτ−q(W+w0)−sK\tau^{-q}(W+w_0)^{-s} term captures the persistent error reduction from warmup
  • w0=0.032w_0 = 0.032 (in thousands of updates) keeps the predictor well-defined at zero warmup

Selection Rule

For warmup selection, only differences between candidate losses matter:

ΔL(W,T;Wref)=A[τW−p−τref−p]+C[τW−q(W+w0Wref+w0)−s−τref−q]\Delta L(W, T; W_{ref}) = A[\tau_W^{-p} - \tau_{ref}^{-p}] + C\left[\tau_W^{-q}\left(\frac{W + w_0}{W_{ref} + w_0}\right)^{-s} - \tau_{ref}^{-q}\right]

Empirical Validation / Results

Key Empirical Findings

1. Peak learning rate determines warmup regime:

RegimePeak LRBehavior
ConservativeLow (e.g., 0.0003)Little or no warmup; preferred duration changes little with horizon
AggressiveHigh (e.g., 0.007)Optimal warmup grows substantially with training horizon

2. Ranking reversal: At high peak rates, a longer warmup can start worse but become better later, so its benefit persists beyond the warmup duration itself.

3. Best runs use long warmups (Table 1):

ModelPeak LRWarmup (k)Fraction of horizon (T=128k)Loss
60M0.0043023.4%3.404
100M0.0055039.1%3.246
350M0.0044031.3%2.983

4. Shortest successful warmup increases with:

  • Higher peak learning rates
  • Smaller batch sizes
  • Larger models

Theoretical Results

Two-mode optimum: With a low-curvature mode (where ρs>gs\rho_s > g_s) and a near-stability mode (where ge>ρeg_e > \rho_e), the optimal warmup is:

Wquad⋆(T,η)=Π[0,T][(ρs−ρe)T+log⁡(Aeae/(Asds))ds+ae]W^\star_{quad}(T, \eta) = \Pi_{[0,T]}\left[\frac{(\rho_s - \rho_e)T + \log(A_e a_e / (A_s d_s))}{d_s + a_e}\right]

If 0<hs<he/20 < h_s < h_e/2 and ηhe<2\eta h_e < 2, the long-horizon optimum changes at ηc=2/(hs+he)\eta_c = 2/(h_s + h_e):

  • Zero for η<ηc\eta < \eta_c
  • Bounded at η=ηc\eta = \eta_c
  • A fraction of T for η>ηc\eta > \eta_c, growing with η\eta

Fit Quality and Prediction

Table 2: Absolute-fit accuracy and selection quality

EvaluationMedian R²RMSEMean regret
All horizons (descriptive)0.9964.922.00 ± 1.51
Alternating holdout0.9964.982.05 ± 1.61

RMSE and regret in 10⁻³ loss units.

Warmup Growth Exponent

The fitted growth exponent β=(p+1−q)/(s+1)\beta = (p+1-q)/(s+1) from the loss law predicts:

W⋆∼(2sKAp)1/(s+1)TβW^\star \sim \left(\frac{2sK}{Ap}\right)^{1/(s+1)} T^{\beta}
  • β≤0\beta \leq 0: zero or bounded warmup
  • 0<β<10 < \beta < 1: sublinear growth
  • β≥1\beta \geq 1: proportional growth

Extrapolation Results

Table 3: Fits through 32k select low-regret warmups at later horizons

SelectorParametersAlternatingFour horizons128k
Absolute fit62.05 ± 1.612.26 ± 0.663.05
Difference fit51.45 ± 2.221.58 ± 0.772.61
1k target015.47 ± 10.8013.79 ± 0.2213.56
10% target013.39 ± 4.1012.58 ± 0.9011.72

Mean regret in 10⁻³ loss units.

Table 4: Three short runs support longer-horizon warmup prediction

Fit horizonAbsoluteDifferenceFixed dur.Fixed frac.1k10%
20k (29 families)2.83 ± 0.164.22 ± 0.299.31 ± 1.2111.72 ± 0.8214.28 ± 0.3013.42 ± 0.96
32k (33 families)2.51 ± 0.272.62 ± 0.367.61 ± 1.128.47 ± 0.6213.79 ± 0.2212.58 ± 0.90

Mean regret over 50k, 75k, 100k, and 128k, in 10⁻³ loss units.

Long-Horizon Extrapolation (up to 1M updates)

Table 5: Long-horizon validation-loss comparison (100M model, batch size 128)

LR = 0.5×10⁻³LR = 2×10⁻³LR = 5×10⁻³
T (k)250500100025050010002505001000
1k-warmup loss3.3703.3373.3113.4333.4213.4023.3713.3473.322
10% loss3.3753.3473.3393.4183.3673.3403.3633.3353.314
Ŵ₅₀0.60.70.847.777.5125.9129.5241.7451.1
Loss3.3723.3363.3113.4123.3623.3383.3593.3253.303
Ŵ₇₅0.81.11.256.695.2160.0116.0213.5392.9
Loss3.3713.3363.3123.4113.3593.3363.3583.3243.303

Horizons and predicted warmups in thousands of updates.

Theoretical and Practical Implications

Theoretical Implications

  1. Unified explanation: The quadratic model explains several familiar properties of warmup through a single tradeoff—between giving up early progress and reducing persistent error near the stability edge.

  2. Stability requirements: The model predicts the minimum warmup fraction to avoid divergence:

WT≥κ(αe)=−ρ(αe)g(αe)−ρ(αe)=αelog⁡(αe−1)αe+log⁡(αe−1)\frac{W}{T} \geq \kappa(\alpha_e) = \frac{-\rho(\alpha_e)}{g(\alpha_e) - \rho(\alpha_e)} = \frac{\alpha_e \log(\alpha_e - 1)}{\alpha_e + \log(\alpha_e - 1)}

which increases with the peak rate and eventually reaches one.

  1. Scaling law framework: Warmup duration joins other optimization hyperparameters (learning rate, batch size, weight decay) that have their own scaling laws with training budget.

Practical Implications

  1. Warmup should be tuned jointly with peak learning rate and training budget—fixing warmup in advance can change which learning rate appears best.

  2. Practical prediction protocol: Fit the loss law using only three short runs (warmups of 2k, 8k, and 16k updates) with checkpoints through 20k or 32k updates, then predict warmup at substantially longer horizons with low regret.

  3. Regime awareness: Practitioners should expect:

    • Low peak rates → short or bounded warmup
    • High peak rates → warmup that grows with the horizon
  4. Robustness: The warmup ordering is largely preserved when decay is appended (warmup-stable-decay schedules), and the behavior holds across different warmup shapes (linear, half-cosine, concave-quadratic) and optimizers (AdamW, Muon).

Conclusion

The paper demonstrates that warmup duration is neither fixed across training budgets nor tied to a fixed fraction of the run. The key findings are:

  1. Regime-dependent scaling: Lower peak learning rates favor short or bounded warmup, while higher rates favor durations that grow with the training horizon.

  2. Mechanistic explanation: A simple quadratic model explains this shift as a balance between sacrificing early progress and reducing error that would otherwise persist later in training.

  3. Compact loss law: The derived loss law L(W,T)=L∞+Aτ−p+Kτ−q(W+w0)−sL(W, T) = L_\infty + A\tau^{-p} + K\tau^{-q}(W + w_0)^{-s} captures the observed warmup regimes and can be fit using only three short trajectories to guide warmup choices at longer horizons.

  4. Practical recommendation: Treat warmup duration as a horizon-dependent hyperparameter, tuned jointly with peak learning rate and training budget, rather than as a fixed training heuristic.

Future Directions

  • Model-scale extrapolation: The authors show a correspondence between model size and effective peak learning rate, suggesting a model-size correction could extend predictions across scales.
  • Understanding how warmup interacts with other schedule components (decay phase, annealing) in more complex schedules.
  • Extending the analysis to other optimizers and training regimes beyond the ones studied.

Related papers