# Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons

> Optimal learning-rate warmup duration scales with training horizon only at high peak learning rates, following a regime-dependent law predictable from three short runs.

- **Source:** [arXiv](https://arxiv.org/abs/2609.33041)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/zPSYzB
- **Whiteboard:** https://picx.dev/p/zPSYzB/image

## Summary

## Summary (Overview)

- This paper investigates how the optimal learning-rate warmup duration scales with the training horizon in language model training, showing that it is neither fixed nor a constant fraction of the horizon but follows a regime-dependent scaling law.
- The authors identify three distinct regimes: essentially no warmup at low peak learning rates, bounded warmup at intermediate rates, and warmup durations that grow (sometimes near-proportionally) with the training horizon at high peak rates.
- A simple quadratic model with two modes (low-curvature and near-stability) explains the tradeoff: warmup sacrifices early progress in well-conditioned directions but removes persistent error in directions near the stability edge, with higher peak rates shifting the balance toward longer warmup.
- The paper derives a compact loss law $$L(W,T) = L_\infty + A\tau^{-p} + K\tau^{-q}(W + w_0)^{-s}$$ that captures the warmup-horizon tradeoff and can be fit using as few as three short training runs to predict optimal warmup at substantially longer horizons.
- The findings suggest treating warmup duration as a horizon-dependent hyperparameter that should be tuned jointly with peak learning rate and training budget, rather than as a fixed heuristic.

## Introduction and Theoretical Foundation

### Background and Motivation

Learning-rate warmup is ubiquitous in language-model training, yet its duration remains largely heuristic. Common approaches use either:
- **Fixed number of warmup steps** (Vaswani et al., 2017; Groeneveld et al., 2024; Kaplan et al., 2020)
- **Fixed fraction of the training horizon** (Black et al., 2022)

These choices imply very different scaling as training gets longer, raising the question: *When should warmup stay fixed, and when should it grow with the horizon?*

### Prior Work

Previous explanations of warmup include:
- Adaptive-optimizer statistics and update magnitudes (Liu et al., 2020; Ma & Yarats, 2021)
- Loss curvature and access to larger subsequent learning rates (Gilmer et al., 2022; Kalra & Barkeshli, 2024)
- Early changes in parameters and representations (Gotmare et al., 2019; Kosson et al., 2024)
- Suboptimality-dependent smoothness (Liu et al., 2025b; Alimisis et al., 2026; Riabinin et al., 2026)

However, explaining *why* warmup helps does not determine its *optimal duration*. That requires balancing the persistent benefit of longer warmup against the progress lost to smaller initial steps.

### Theoretical Foundation: Quadratic Model

The paper uses a quadratic model as a standard lens for studying optimization effects:

$$f(\theta) - f^\star = \frac{1}{2}(\theta - \theta^\star)^\top H(\theta - \theta^\star), \quad H \succeq 0$$

In a direction of curvature $h > 0$, one update with learning rate $\eta_t$ multiplies that direction's loss by $|1 - \eta_t h|^2$. Defining $\alpha = \eta h$, the contraction rate at the peak learning rate and its average over linear warmup are:

$$\rho(\alpha) = -2\log|1-\alpha|, \quad g(\alpha) = \int_0^1 \rho(\alpha u)\,du$$

The loss remaining after warmup and the subsequent peak-rate phase is:

$$M_{T,W}(\alpha) \approx e^{-g(\alpha)W - \rho(\alpha)(T-W)}$$

## Methodology

### Experimental Design

- **Models**: Llama-style language models of sizes 60M, 100M, and 350M parameters
- **Training**: Pretrained on C4 English with t5-base SentencePiece tokenizer, sequence length 256, batch size 512 (131,072 tokens per update)
- **Schedule**: Linear warmup followed by constant peak rate (warmup–stable), which allows one trajectory to be evaluated at many horizons
- **Grid**: 33 model-size and peak-rate families, with warmup ranges from 0 to 64,000 updates
- **Peak learning rates**: 0.1 to 8 × 10⁻³

### The Loss Law

The paper derives a compact loss law over warmup duration $W$ and training horizon $T$:

$$L(W, T) = L_\infty + A\tau^{-p} + K\tau^{-q}(W + w_0)^{-s}$$

where:
- $\tau = T - W/2$ is the effective peak-rate-equivalent training amount (a warmup step counts as half a peak-rate step)
- $L_\infty$ is the irreducible loss
- The $A\tau^{-p}$ term captures the **cost of delayed progress** (longer warmup reduces $\tau$)
- The $K\tau^{-q}(W+w_0)^{-s}$ term captures the **persistent error reduction** from warmup
- $w_0 = 0.032$ (in thousands of updates) keeps the predictor well-defined at zero warmup

### Selection Rule

For warmup selection, only differences between candidate losses matter:

$$\Delta L(W, T; W_{ref}) = A[\tau_W^{-p} - \tau_{ref}^{-p}] + C\left[\tau_W^{-q}\left(\frac{W + w_0}{W_{ref} + w_0}\right)^{-s} - \tau_{ref}^{-q}\right]$$

## Empirical Validation / Results

### Key Empirical Findings

**1. Peak learning rate determines warmup regime:**

| Regime | Peak LR | Behavior |
|--------|---------|----------|
| Conservative | Low (e.g., 0.0003) | Little or no warmup; preferred duration changes little with horizon |
| Aggressive | High (e.g., 0.007) | Optimal warmup grows substantially with training horizon |

**2. Ranking reversal**: At high peak rates, a longer warmup can start worse but become better later, so its benefit persists beyond the warmup duration itself.

**3. Best runs use long warmups** (Table 1):

| Model | Peak LR | Warmup (k) | Fraction of horizon (T=128k) | Loss |
|-------|---------|------------|------------------------------|------|
| 60M | 0.004 | 30 | 23.4% | 3.404 |
| 100M | 0.005 | 50 | 39.1% | 3.246 |
| 350M | 0.004 | 40 | 31.3% | 2.983 |

**4. Shortest successful warmup** increases with:
- Higher peak learning rates
- Smaller batch sizes
- Larger models

### Theoretical Results

**Two-mode optimum**: With a low-curvature mode (where $\rho_s > g_s$) and a near-stability mode (where $g_e > \rho_e$), the optimal warmup is:

$$W^\star_{quad}(T, \eta) = \Pi_{[0,T]}\left[\frac{(\rho_s - \rho_e)T + \log(A_e a_e / (A_s d_s))}{d_s + a_e}\right]$$

If $0 < h_s < h_e/2$ and $\eta h_e < 2$, the long-horizon optimum changes at $\eta_c = 2/(h_s + h_e)$:
- **Zero** for $\eta < \eta_c$
- **Bounded** at $\eta = \eta_c$
- **A fraction of T** for $\eta > \eta_c$, growing with $\eta$

### Fit Quality and Prediction

**Table 2: Absolute-fit accuracy and selection quality**

| Evaluation | Median R² | RMSE | Mean regret |
|-----------|-----------|------|-------------|
| All horizons (descriptive) | 0.996 | 4.92 | 2.00 ± 1.51 |
| Alternating holdout | 0.996 | 4.98 | 2.05 ± 1.61 |

*RMSE and regret in 10⁻³ loss units.*

### Warmup Growth Exponent

The fitted growth exponent $\beta = (p+1-q)/(s+1)$ from the loss law predicts:

$$W^\star \sim \left(\frac{2sK}{Ap}\right)^{1/(s+1)} T^{\beta}$$

- $\beta \leq 0$: zero or bounded warmup
- $0 < \beta < 1$: sublinear growth
- $\beta \geq 1$: proportional growth

### Extrapolation Results

**Table 3: Fits through 32k select low-regret warmups at later horizons**

| Selector | Parameters | Alternating | Four horizons | 128k |
|----------|-----------|-------------|---------------|------|
| Absolute fit | 6 | 2.05 ± 1.61 | 2.26 ± 0.66 | 3.05 |
| Difference fit | 5 | 1.45 ± 2.22 | 1.58 ± 0.77 | 2.61 |
| 1k target | 0 | 15.47 ± 10.80 | 13.79 ± 0.22 | 13.56 |
| 10% target | 0 | 13.39 ± 4.10 | 12.58 ± 0.90 | 11.72 |

*Mean regret in 10⁻³ loss units.*

**Table 4: Three short runs support longer-horizon warmup prediction**

| Fit horizon | Absolute | Difference | Fixed dur. | Fixed frac. | 1k | 10% |
|------------|----------|------------|------------|-------------|-----|-----|
| 20k (29 families) | 2.83 ± 0.16 | 4.22 ± 0.29 | 9.31 ± 1.21 | 11.72 ± 0.82 | 14.28 ± 0.30 | 13.42 ± 0.96 |
| 32k (33 families) | 2.51 ± 0.27 | 2.62 ± 0.36 | 7.61 ± 1.12 | 8.47 ± 0.62 | 13.79 ± 0.22 | 12.58 ± 0.90 |

*Mean regret over 50k, 75k, 100k, and 128k, in 10⁻³ loss units.*

### Long-Horizon Extrapolation (up to 1M updates)

**Table 5: Long-horizon validation-loss comparison** (100M model, batch size 128)

| | LR = 0.5×10⁻³ | | | LR = 2×10⁻³ | | | LR = 5×10⁻³ | | |
|---|---|---|---|---|---|---|---|---|---|
| T (k) | 250 | 500 | 1000 | 250 | 500 | 1000 | 250 | 500 | 1000 |
| 1k-warmup loss | 3.370 | 3.337 | 3.311 | 3.433 | 3.421 | 3.402 | 3.371 | 3.347 | 3.322 |
| 10% loss | 3.375 | 3.347 | 3.339 | 3.418 | 3.367 | 3.340 | 3.363 | 3.335 | 3.314 |
| Ŵ₅₀ | 0.6 | 0.7 | 0.8 | 47.7 | 77.5 | 125.9 | 129.5 | 241.7 | 451.1 |
| Loss | 3.372 | 3.336 | 3.311 | 3.412 | 3.362 | 3.338 | 3.359 | 3.325 | 3.303 |
| Ŵ₇₅ | 0.8 | 1.1 | 1.2 | 56.6 | 95.2 | 160.0 | 116.0 | 213.5 | 392.9 |
| Loss | 3.371 | 3.336 | 3.312 | 3.411 | 3.359 | 3.336 | 3.358 | 3.324 | 3.303 |

*Horizons and predicted warmups in thousands of updates.*

## Theoretical and Practical Implications

### Theoretical Implications

1. **Unified explanation**: The quadratic model explains several familiar properties of warmup through a single tradeoff—between giving up early progress and reducing persistent error near the stability edge.

2. **Stability requirements**: The model predicts the minimum warmup fraction to avoid divergence:
$$\frac{W}{T} \geq \kappa(\alpha_e) = \frac{-\rho(\alpha_e)}{g(\alpha_e) - \rho(\alpha_e)} = \frac{\alpha_e \log(\alpha_e - 1)}{\alpha_e + \log(\alpha_e - 1)}$$
which increases with the peak rate and eventually reaches one.

3. **Scaling law framework**: Warmup duration joins other optimization hyperparameters (learning rate, batch size, weight decay) that have their own scaling laws with training budget.

### Practical Implications

1. **Warmup should be tuned jointly** with peak learning rate and training budget—fixing warmup in advance can change which learning rate appears best.

2. **Practical prediction protocol**: Fit the loss law using only **three short runs** (warmups of 2k, 8k, and 16k updates) with checkpoints through 20k or 32k updates, then predict warmup at substantially longer horizons with low regret.

3. **Regime awareness**: Practitioners should expect:
   - Low peak rates → short or bounded warmup
   - High peak rates → warmup that grows with the horizon

4. **Robustness**: The warmup ordering is largely preserved when decay is appended (warmup-stable-decay schedules), and the behavior holds across different warmup shapes (linear, half-cosine, concave-quadratic) and optimizers (AdamW, Muon).

## Conclusion

The paper demonstrates that **warmup duration is neither fixed across training budgets nor tied to a fixed fraction of the run**. The key findings are:

1. **Regime-dependent scaling**: Lower peak learning rates favor short or bounded warmup, while higher rates favor durations that grow with the training horizon.

2. **Mechanistic explanation**: A simple quadratic model explains this shift as a balance between sacrificing early progress and reducing error that would otherwise persist later in training.

3. **Compact loss law**: The derived loss law $$L(W, T) = L_\infty + A\tau^{-p} + K\tau^{-q}(W + w_0)^{-s}$$ captures the observed warmup regimes and can be fit using only three short trajectories to guide warmup choices at longer horizons.

4. **Practical recommendation**: Treat warmup duration as a **horizon-dependent hyperparameter**, tuned jointly with peak learning rate and training budget, rather than as a fixed training heuristic.

### Future Directions

- Model-scale extrapolation: The authors show a correspondence between model size and effective peak learning rate, suggesting a model-size correction could extend predictions across scales.
- Understanding how warmup interacts with other schedule components (decay phase, annealing) in more complex schedules.
- Extending the analysis to other optimizers and training regimes beyond the ones studied.

---

_Markdown view of https://picx.dev/p/zPSYzB, served by PicX — AI-generated visual whiteboard summaries of research papers._
