Summary (Overview)
- Core contribution: This paper provides a rigorous theory of when and how power-law learning curves emerge in noisy online SGD with linear random features, and how joint learning-rate/batch-size schedules transform these scaling laws.
- Key mechanism: The loss is decomposed into a forcing term (propagating unresolved target error) and a memory kernel (propagating stochastic-error injections), governed by an exact Volterra equation. Power laws in either component arise if and only if the cumulative weighted spectral mass near zero has the corresponding scaling—not from coordinatewise power laws of individual eigenvalues or target coefficients.
- 3+3(+2) regime map: In the power-law random-feature (PLRF) model, the authors derive a phase diagram with 3 long-memory (LM), 3 integrable-memory (IM), and 2 finite-bulk (FB) propagation regimes, each with distinct forcing and noise response asymptotics and phase-dependent optimal compute rates.
- Schedule transformation theorem: A sharp classification of when a joint schedule preserves, changes, or destroys the clean power law, including a "memory ceiling" beyond which reducing late-stage noise cannot improve decay.
- LLM validation: Controlled 300M nanoGPT experiments show that (1) schedules with matched paths nearly coincide in intrinsic time, (2) a theory-derived forcing-memory surrogate fitted on one schedule predicts held-out schedules without refitting, and (3) fitted exponents across OpenWebText, FineWeb, and peS2o place LLM responses near , the LM/IM boundary.
Introduction and Theoretical Foundation
Background and Motivation
Power-law learning curves are widely used in LLM pre-training to forecast progress and allocate compute (Hestness et al., 2017; Kaplan et al., 2020). However, the paper argues that a scaling exponent is not a fixed property of a model and its data—changing learning-rate or batch-size schedules can alter both optimization progress and stochastic error accumulation.
The authors distinguish two problems:
- Origin: When do underlying learning dynamics produce power-law response components?
- Transfer: Once such components exist, when does a joint schedule preserve, change, or destroy their law in the observed loss?
Theoretical Foundation
The paper studies a linear random-feature model trained by noisy online SGD (Rahimi and Recht, 2007; Mei and Montanari, 2022). The data generation is:
with Gaussian data , .
The model is with fixed random features (i.i.d. entries), trained by SGD with mini-batch size and learning rate :
Here is the intrinsic time (optimization clock) and is the variance injected per unit intrinsic time.
Methodology
Exact Volterra Equation
Conditional on the representation, the excess risk satisfies an exact recursion:
where the forcing term and memory kernel are:
with the one-step survival factor:
Spectral Criterion
The key theoretical result is the componentwise spectral criterion:
Theorem 3.1 (informal): With :
where the cumulative weighted spectral masses are:
PLRF Regime Map
For the canonical power-law setting:
the moving cutoff gives:
The phase map:
- LM (long memory): (non-integrable survival kernel)
- IM (integrable memory):
- FB (finite bulk): (no width-independent exists)
Schedule Transformation
For a varying schedule, the noisy-clean gap satisfies:
Theorem 5.1 gives the noise exponent for and :
Empirical Validation / Results
Finite-Width Simulations
Figure 2 validates Theorem 3.1 with , , :
- Forcing and cumulative forcing mass both decay faster than every inverse power of
- Memory response and memory mass both scale as
- The minibatch-SGD loss follows the exponentially decaying forcing initially, then crosses over to the memory tail
nanoGPT Experiments (300M)
Test 1 — Factorization collapse: Schedules with matched paths (WSD and 8-1-1) follow different trajectories in optimizer step but nearly coincide in intrinsic time (Figures 1a, 1b).
Test 2 — Zero-refit transfer: A 7-parameter surrogate fitted only to the 8-1-1 trajectory predicts the held-out WSD trajectory without refitting (Figure 5):
Test 3 — Regime identification: Fitted exponents across datasets:
| Dataset | Fitted |
|---|---|
| OpenWebText | 1.017 |
| FineWeb | 0.952 |
| peS2o V2 | 0.995 |
These values cluster near , placing LLM responses near the LM/IM boundary, corresponding to effective (where the 4+3 theory predicts the near-square-root compute-optimal width scaling reported by Chinchilla).
Theoretical and Practical Implications
Theoretical Implications
-
Power laws are dynamical responses, not static properties: The paper proves that power-law learning curves emerge from cumulative weighted spectral mass near zero, not from coordinatewise power laws. This means:
- Power-law data are neither sufficient nor necessary for power-law responses
- Irregular or oscillatory targets can produce canonical exponents
- A finite spectrum leads to exponential decay (possibly after a long intermediate scaling window)
-
Schedule dependence is fundamental: The loss curve is jointly shaped by spectrum, target alignment, noise, and schedule. The observed loss is not determined by componentwise power laws separately but by the competition between clean target-error decay and accumulated stochastic response.
-
Non-identifiability: The observed loss identifies the ratio but not learning rate and batch size separately below the memory ceiling.
Practical Implications
- Schedule design: The optimal ratio path satisfies:
allocating more samples where injections are both large and likely to survive until time .
- Optimal resource rates (Table 1, with ):
| Regime and branch | ||
|---|---|---|
- Cross-schedule prediction: The forcing-memory surrogate enables predicting held-out schedules without refitting, providing a practical tool for LLM training planning.
Conclusion
The paper establishes that power-law learning curves are dynamical responses jointly shaped by spectrum, target alignment, noise, and schedule. Key takeaways:
-
Target-weighted and squared-spectrum masses determine forcing and memory respectively, with sharp if-and-only-if criteria for component power laws.
-
The joint schedule with and controls how memory accumulates, with a sharp preserve-change-destroy classification.
-
In the canonical PLRF model, the map classifies propagation regimes with phase-dependent optimal rates.
-
LLM experiments confirm the proxy's predictive power: the surrogate transfers across schedules without refitting, and the fitted (effective ) is robust across datasets and model scales.
Future directions include extending the analysis to optimizers beyond plain SGD (the paper notes Muon trajectories collapse under rather than ), and further exploring the connection to compute-optimal scaling laws like Chinchilla.
Related papers
- The Winner's Curse in LLM Self-Improvement Loops: Selection Noise, Lock-in, and Acceptance Rules
Self-improving LLM loops on small reused evaluation sets inflate reported gains by 13–20 points, with most proposals after the first rewrite being harmful.
- Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
Verifier evolution lets agents self-improve without ground truth, but only anchor discipline—not detector lifecycle—prevents collapse into vacuous always-pass grading.
- Stateless Language Agents: Scaling Long-Horizon Automated Research
Stateless Language Agents, where the harness owns all research state and reconstructs fresh contexts per invocation, outperform stateful agent frameworks on long-horizon tasks, reaching baseline final performance with over 84% fewer tokens.