# From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes

> Power-law learning curves emerge from cumulative weighted spectral mass near zero, not individual eigenvalues, and are jointly shaped by spectrum, target, noise, and schedule.

- **Source:** [arXiv](https://arxiv.org/abs/2609.40148)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/B0bfZx
- **Whiteboard:** https://picx.dev/p/B0bfZx/image

## Summary

## Summary (Overview)

- **Core contribution**: This paper provides a rigorous theory of when and how power-law learning curves emerge in noisy online SGD with linear random features, and how joint learning-rate/batch-size schedules transform these scaling laws.
- **Key mechanism**: The loss is decomposed into a **forcing term** (propagating unresolved target error) and a **memory kernel** (propagating stochastic-error injections), governed by an exact Volterra equation. Power laws in either component arise if and only if the *cumulative weighted spectral mass* near zero has the corresponding scaling—not from coordinatewise power laws of individual eigenvalues or target coefficients.
- **3+3(+2) regime map**: In the power-law random-feature (PLRF) model, the authors derive a phase diagram with 3 long-memory (LM), 3 integrable-memory (IM), and 2 finite-bulk (FB) propagation regimes, each with distinct forcing and noise response asymptotics and phase-dependent optimal compute rates.
- **Schedule transformation theorem**: A sharp classification of when a joint schedule $r_t = B_t/\eta_t$ preserves, changes, or destroys the clean power law, including a "memory ceiling" beyond which reducing late-stage noise cannot improve decay.
- **LLM validation**: Controlled 300M nanoGPT experiments show that (1) schedules with matched $B/\eta$ paths nearly coincide in intrinsic time, (2) a theory-derived forcing-memory surrogate fitted on one schedule predicts held-out schedules without refitting, and (3) fitted exponents across OpenWebText, FineWeb, and peS2o place LLM responses near $q_\mathcal{K} \approx 1$, the LM/IM boundary.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Power-law learning curves are widely used in LLM pre-training to forecast progress and allocate compute (Hestness et al., 2017; Kaplan et al., 2020). However, the paper argues that a scaling exponent is **not** a fixed property of a model and its data—changing learning-rate or batch-size schedules can alter both optimization progress and stochastic error accumulation.

The authors distinguish two problems:
1. **Origin**: When do underlying learning dynamics produce power-law response components?
2. **Transfer**: Once such components exist, when does a joint schedule preserve, change, or destroy their law in the observed loss?

### Theoretical Foundation

The paper studies a **linear random-feature model** trained by noisy online SGD (Rahimi and Recht, 2007; Mei and Montanari, 2022). The data generation is:

$$y = f_\star(\mathbf{x}) + \varepsilon, \qquad f_\star(\mathbf{x}) := \langle \mathbf{x}, \boldsymbol{\theta}^\star \rangle, \qquad \mathbb{E}[\varepsilon] = 0, \qquad \mathbb{E}[\varepsilon^2] = \sigma^2$$

with Gaussian data $\mathbf{x} = \boldsymbol{\Lambda}^{1/2}\mathbf{z} \in \mathbb{R}^d$, $\boldsymbol{\Lambda} = \mathrm{diag}(\lambda_1, \dots, \lambda_d) \succ \mathbf{0}$.

The model is $f_a(x) = \langle W^\top x, a \rangle$ with fixed random features $W \in \mathbb{R}^{d \times m}$ (i.i.d. $\mathcal{N}(0, 1/m)$ entries), trained by SGD with mini-batch size $B_t$ and learning rate $\eta_t$:

$$\boldsymbol{a}_{t+1} = \boldsymbol{a}_t - \frac{\eta_t}{B_t} \sum_{i=1}^{B_t} \boldsymbol{W}^\top \boldsymbol{x}_t^i \left(f_{\boldsymbol{a}_t}(\boldsymbol{x}_t^i) - y_t^i\right), \quad T_t := \sum_{s < t} \eta_s, \quad r_t := \frac{B_t}{\eta_t} \tag{2.1}$$

Here $T_t$ is the **intrinsic time** (optimization clock) and $1/r_t$ is the variance injected per unit intrinsic time.

---

## Methodology

### Exact Volterra Equation

Conditional on the representation, the excess risk satisfies an exact recursion:

$$R_{\sigma,t} = F_{\boldsymbol{W}}(t) + \sum_{s=0}^{t-1} K_{\boldsymbol{W}}(t-1-s)(R_{\sigma,s} + \sigma^2) \tag{2.3}$$

where the **forcing term** and **memory kernel** are:

$$F_{\boldsymbol{W}}(t) := \sum_j |\langle \hat{\boldsymbol{u}}_j, \boldsymbol{\Lambda}^{1/2}\boldsymbol{\theta}^\star\rangle|^2 q_\eta(\hat{\lambda}_j)^t, \qquad K_{\boldsymbol{W}}(t) := \frac{\eta^2}{B} \sum_j \hat{\lambda}_j^2 q_\eta(\hat{\lambda}_j)^t$$

with the one-step survival factor:

$$q_\eta(\lambda) := 1 - 2\eta\lambda + (1 + 1/B)\eta^2\lambda^2 \tag{2.2}$$

### Spectral Criterion

The key theoretical result is the **componentwise spectral criterion**:

**Theorem 3.1 (informal)**: With $x = T^{-1} \downarrow 0$:

$$\nu_{\boldsymbol{W}}^{\mathcal{F}}((0,x]) \propto x^{q_{\mathcal{F}}} \Longleftrightarrow F_{\boldsymbol{W},>0}(t) \propto T^{-q_{\mathcal{F}}}, \quad \nu_{\boldsymbol{W}}^{\mathcal{K}}((0,x]) \propto x^{q_{\mathcal{K}}} \Longleftrightarrow \frac{B}{\eta^2}K_{\boldsymbol{W}}(t) \propto T^{-q_{\mathcal{K}}}$$

where the cumulative weighted spectral masses are:

$$\nu_{\boldsymbol{W}}^{\mathcal{F}}((0,x]) = \text{target energy in learnable directions slower than } x^{-1}, \quad \nu_{\boldsymbol{W}}^{\mathcal{K}}((0,x]) = \text{squared spectral mass below } x^{-1}$$

### PLRF Regime Map

For the canonical power-law setting:

$$\lambda_j = j^{-2\alpha}, \qquad \theta_j^\star = j^{-\beta}, \qquad 2\alpha + 2\beta > 1$$

the moving cutoff $\lambda_{j_T} \asymp T^{-1}$ gives:

$$F(T) \propto T^{-q_{\mathcal{F}}}, \quad q_{\mathcal{F}} := \frac{2(\alpha+\beta)-1}{2\alpha}, \qquad \frac{B}{\eta^2}K(T) \propto T^{-q_{\mathcal{K}}}, \quad q_{\mathcal{K}} := 2 - \frac{1}{2\alpha}$$

The phase map:
- **LM** (long memory): $0 < q_{\mathcal{K}} < 1$ (non-integrable survival kernel)
- **IM** (integrable memory): $q_{\mathcal{K}} > 1$
- **FB** (finite bulk): $0 < \alpha < 1/4$ (no width-independent $q_{\mathcal{K}}$ exists)

### Schedule Transformation

For a varying schedule, the noisy-clean gap satisfies:

$$R_{\sigma,t} - R_{0,t} = \sum_{s<t} K_{t,s}\left[\sigma^2 + R_{\sigma,s} - R_{0,s}\right], \quad \sum_{s<t} K_{t,s} \asymp \int_0^{T_t} \frac{k(T_t - u)}{r(u)}\mathrm{d}u$$

**Theorem 5.1** gives the noise exponent for $k(T) \sim c_\mathcal{K} T^{-q_{\mathcal{K}}}$ and $r(T) \sim c_r T^{\vartheta}$:

$$q_{\mathcal{N}}(\vartheta) := \begin{cases} \vartheta + q_{\mathcal{K}} - 1, & \text{LM}, \quad 1 - q_{\mathcal{K}} < \vartheta < 1, \\ q_{\mathcal{K}}, & \text{LM}, \quad \vartheta > 1, \\ \min\{\vartheta, q_{\mathcal{K}}\}, & \text{IM}, \quad \vartheta > 0. \end{cases}$$

---

## Empirical Validation / Results

### Finite-Width Simulations

Figure 2 validates Theorem 3.1 with $\lambda_j = j^{-0.8}$, $|\theta_j^\star|^2 = e^{-j}$, $d = 524,288$:
- Forcing and cumulative forcing mass both decay faster than every inverse power of $T$
- Memory response and memory mass both scale as $T^{-3/4}$
- The minibatch-SGD loss follows the exponentially decaying forcing initially, then crosses over to the $T^{-3/4}$ memory tail

### nanoGPT Experiments (300M)

**Test 1 — Factorization collapse**: Schedules with matched $B_t/\eta_t$ paths (WSD and 8-1-1) follow different trajectories in optimizer step but nearly coincide in intrinsic time (Figures 1a, 1b).

**Test 2 — Zero-refit transfer**: A 7-parameter surrogate fitted only to the 8-1-1 trajectory predicts the held-out WSD trajectory without refitting (Figure 5):

$$\widehat{L}(T) = L_\infty + A_{\mathcal{F}}(1+T)^{-q_{\mathcal{F}}} + \int_0^T \frac{A_0 + A_1(1+u)^{-q_{\mathcal{F}}}}{r(u)}\left(1 + c_{\mathcal{K}}(T-u)\right)^{-q_{\mathcal{K}}}\mathrm{d}u \tag{5.3}$$

**Test 3 — Regime identification**: Fitted exponents across datasets:

| Dataset | Fitted $q_{\mathcal{K}}$ |
|---------|------------------------|
| OpenWebText | 1.017 |
| FineWeb | 0.952 |
| peS2o V2 | 0.995 |

These values cluster near $q_{\mathcal{K}} \approx 1$, placing LLM responses near the LM/IM boundary, corresponding to effective $\alpha \approx 1/2$ (where the 4+3 theory predicts the near-square-root compute-optimal width scaling reported by Chinchilla).

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **Power laws are dynamical responses, not static properties**: The paper proves that power-law learning curves emerge from cumulative weighted spectral mass near zero, not from coordinatewise power laws. This means:
   - Power-law data are neither sufficient nor necessary for power-law responses
   - Irregular or oscillatory targets can produce canonical exponents
   - A finite spectrum leads to exponential decay (possibly after a long intermediate scaling window)

2. **Schedule dependence is fundamental**: The loss curve is jointly shaped by spectrum, target alignment, noise, and schedule. The observed loss is not determined by componentwise power laws separately but by the competition between clean target-error decay and accumulated stochastic response.

3. **Non-identifiability**: The observed loss identifies the ratio $r = B/\eta$ but not learning rate and batch size separately below the memory ceiling.

### Practical Implications

1. **Schedule design**: The optimal ratio path satisfies:

$$r^\star(u) \propto \sqrt{(1+T-u)^{-q_{\mathcal{K}}}[\mathcal{F}(u,m) + \sigma^2]}$$

allocating more samples where injections are both large and likely to survive until time $T$.

2. **Optimal resource rates** (Table 1, with $p = 2\alpha + 2\beta - 1$):

| Regime and branch | $\mathcal{R}_\sigma^\star(D)$ | $\mathcal{R}_\sigma^\star(\mathfrak{f})$ |
|---|---|---|
| $IM_1, \beta < 0$ | $D^{-p/(2\alpha)}$ | $\mathfrak{f}^{-p/(1+2\alpha)}$ |
| $LM_{1,2}; IM_1, \beta > 0$ | $D^{-p/(1+p)}$ | $\mathfrak{f}^{-p/(2+p)}$ |
| $LM_3$ | $D^{-p/(1+p)}$ | $\mathfrak{f}^{-2\alpha p/[p+2\alpha(1+p)]}$ |
| $IM_{2,3}, 1/2 < \alpha \leq 1$ | $D^{-p/(1+p)}$ | $\mathfrak{f}^{-p/(1+p+2\beta)}$ |
| $IM_{2,3}, \alpha > 1$ | $D^{-p/(1+p)}$ | $\mathfrak{f}^{-\alpha/(1+\alpha)}$ |
| $FB_1$ | $D^{-p/(1+p)}$ | $\mathfrak{f}^{-p/(2+p)}$ |
| $FB_2$ | $D^{-2\alpha p/[p(1-2\alpha)+8\alpha^2]}$ | $\mathfrak{f}^{-2\alpha p/[p(2-2\alpha)+8\alpha^2]}$ |

3. **Cross-schedule prediction**: The forcing-memory surrogate enables predicting held-out schedules without refitting, providing a practical tool for LLM training planning.

---

## Conclusion

The paper establishes that power-law learning curves are **dynamical responses jointly shaped by spectrum, target alignment, noise, and schedule**. Key takeaways:

1. Target-weighted and squared-spectrum masses determine forcing and memory respectively, with sharp if-and-only-if criteria for component power laws.

2. The joint schedule $(T_t, r_t)$ with $T_t = \sum_{s<t}\eta_s$ and $r_t = B_t/\eta_t$ controls how memory accumulates, with a sharp preserve-change-destroy classification.

3. In the canonical PLRF model, the $3+3(+2)$ map classifies propagation regimes with phase-dependent optimal rates.

4. LLM experiments confirm the proxy's predictive power: the surrogate transfers across schedules without refitting, and the fitted $q_{\mathcal{K}} \approx 1$ (effective $\alpha \approx 1/2$) is robust across datasets and model scales.

**Future directions** include extending the analysis to optimizers beyond plain SGD (the paper notes Muon trajectories collapse under $B_t/\eta_t^2$ rather than $B_t/\eta_t$), and further exploring the connection to compute-optimal scaling laws like Chinchilla.

---

_Markdown view of https://picx.dev/p/B0bfZx, served by PicX — AI-generated visual whiteboard summaries of research papers._
