Full text not available for this paper
Summary (Overview)
-
Core contribution: This paper proves fast learning-rate transfer in shallow linear networks when the training horizon grows jointly with width, specifically when , under spectral assumptions on the data Gram matrix.
-
Theoretical framework: The authors leverage the notion of "fast hyperparameter transfer" (Ghosh et al., 2026) and show that convergence of the optimal learning rate alone is insufficient; what matters is the rate of convergence relative to the loss convergence rate.
-
Key rate characterization: The transfer rates are decomposed into three horizon-dependent quantities—loss-value sensitivity (), loss-gradient sensitivity (), and local curvature ()—combined with the finite-width perturbation scale .
-
Distributional results: The paper derives limiting distributions for both the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix.
-
Counterexample: A fixed-horizon counterexample (single positive eigenvalue) shows fast transfer can fail even when the hyperparameter gap converges at the standard rate, demonstrating that nondegenerate first-order loss sensitivity is essential.
Introduction and Theoretical Foundation
Background
The paper addresses the growing cost of hyperparameter tuning for large neural networks. Hyperparameter transfer reduces this cost by reusing hyperparameters tuned on small proxy models when training larger models. The framework builds on:
- μP (Maximal Update Parameterization) (Yang and Hu, 2021): Specifies width-dependent initialization and learning-rate scalings that preserve nontrivial feature learning in the infinite-width limit.
- μTransfer (Yang et al., 2021): Tune hyperparameters on a small proxy and transfer to larger models without retuning.
- Fast hyperparameter transfer (Ghosh et al., 2026): Formalizes when transfer is effective—what matters is whether the optimal hyperparameter converges sufficiently fast relative to the rate at which model performance converges.
Key Challenge: Growing Training Horizons
Recent language-model scaling practices motivate increasing training tokens alongside model size, which often increases gradient updates. The joint scaling limit of width and training horizon is substantially more delicate than the fixed-horizon setting, since the limiting objective itself changes as grows.
Model Setup
The model is a linear network with a single trainable hidden matrix:
where , , and . Initialization follows μP scalings: , , and .
Key definitions: (data Gram matrix), with .
Methodology
Exact Loss Dynamics
The key reduction (Proposition 2.1) shows that training with full-batch gradient descent is equivalent to a scalar problem:
where and is the residual.
Infinite-Width Reference
The infinite-width objective decomposes spectrally:
where are distinct positive eigenvalues of , , and are spectral projectors.
Transfer Quantities (Definition 2.2)
- Loss gap:
- Hyperparameter gap:
- Transfer suboptimality gap:
Fast transfer holds when .
Key Assumptions
Assumption 3.2 (Endpoint activity): The largest and smallest positive eigenvalues and of are simple and satisfy and .
Empirical Validation / Results
Theorem 3.1: Long-Horizon Optimal Learning Rate
As , the infinite-width optimal learning rate converges to:
where and . This is related to the classical optimal step size for Richardson iteration.
Lemma 3.3: Joint Initialization Central Limit Theorem
where with , a symmetric Gaussian matrix, and .
Theorem 3.4: Growing-Horizon Fluctuations and Transfer Rates
Under and , with Assumption 3.2:
The resulting rates are:
| Quantity | Rate |
|---|---|
Corollary 3.5: Fast Transfer at Growing Horizons
when .
Theorem 3.6: Fixed-Horizon Rates
For fixed , under nondegeneracy (): , , .
Proposition 3.8: Counterexample to Fast Transfer
With a single positive eigenvalue ():
Fast transfer fails because the loss gap and transfer penalty scale identically.
Theoretical and Practical Implications
Local Perturbation Formulation (Section 4.2)
The paper abstracts the proof structure into a general framework with three scales:
- : Loss-value response to finite-width perturbations
- : Loss-gradient response to finite-width perturbations
- : Local curvature at the infinite-width optimizer
The transfer penalty satisfies:
Fast transfer follows when .
Key Insights
-
Spectral dependence: Different spectra lead to different transfer rates—contrasting with Wen et al. (2026), where rates are more uniform.
-
Directional misalignment: The loss-value and loss-gradient responses to the same perturbation need not be aligned. Some perturbation directions change the optimized loss at leading order while producing no leading-order shift in the optimal learning rate.
-
Stability vs. optimality: The largest eigenvalue determines the stability boundary (), but the long-horizon optimum depends on both endpoints: .
-
Low-rank concentration is insufficient: The top-1 structure of the gradient update holds in both fast-transfer and failure regimes, showing that structural properties alone don't determine transfer quality.
Conclusion
Main takeaways:
- Fast learning-rate transfer persists at growing horizons when under endpoint activity assumptions
- Transfer rates decompose cleanly into width-scale () and horizon-dependent sensitivity scales (, , )
- Nondegenerate first-order loss sensitivity is essential for fast transfer
Limitations and future directions:
- Restricted to linear networks with a single trainable layer
- Extension to nonlinear/deep networks is an important next step
- Current framework relies on nondegenerate first-order sensitivities; developing higher-order analysis would help understand failure cases like Proposition 3.8
- Studying transfer across multiple scaling axes (width and depth) simultaneously remains open
Related papers
- Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion
Local attention-output reconstruction gains do not guarantee final-model fidelity, as residual completion can improve local error while worsening dense-model KL divergence.
- Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training
NormPre separates LLM update normalization from spectral preconditioning, consistently beating AdamW, Muon, and MANO in pretraining while cutting optimizer latency by up to 67%.
- ScAn-Bench: Evaluating Scaling Analysis Methodology
ScAn-Bench introduces the first surrogate benchmarks for scaling analysis, revealing no universally optimal data acquisition and extrapolation strategy exists across LLMs and VLMs.