Full text not available for this paper

Summary (Overview)

  • Core contribution: This paper proves fast learning-rate transfer in shallow linear networks when the training horizon grows jointly with width, specifically when T=o(n)T = o(\sqrt{n}), under spectral assumptions on the data Gram matrix.

  • Theoretical framework: The authors leverage the notion of "fast hyperparameter transfer" (Ghosh et al., 2026) and show that convergence of the optimal learning rate alone is insufficient; what matters is the rate of convergence relative to the loss convergence rate.

  • Key rate characterization: The transfer rates are decomposed into three horizon-dependent quantities—loss-value sensitivity (ρL,T\rho_{L,T}), loss-gradient sensitivity (ρG,T\rho_{G,T}), and local curvature (κT\kappa_T)—combined with the finite-width perturbation scale ϵn=n−1/2\epsilon_n = n^{-1/2}.

  • Distributional results: The paper derives limiting distributions for both the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix.

  • Counterexample: A fixed-horizon counterexample (single positive eigenvalue) shows fast transfer can fail even when the hyperparameter gap converges at the standard n−1/2n^{-1/2} rate, demonstrating that nondegenerate first-order loss sensitivity is essential.

Introduction and Theoretical Foundation

Background

The paper addresses the growing cost of hyperparameter tuning for large neural networks. Hyperparameter transfer reduces this cost by reusing hyperparameters tuned on small proxy models when training larger models. The framework builds on:

  • μP (Maximal Update Parameterization) (Yang and Hu, 2021): Specifies width-dependent initialization and learning-rate scalings that preserve nontrivial feature learning in the infinite-width limit.
  • μTransfer (Yang et al., 2021): Tune hyperparameters on a small proxy and transfer to larger models without retuning.
  • Fast hyperparameter transfer (Ghosh et al., 2026): Formalizes when transfer is effective—what matters is whether the optimal hyperparameter converges sufficiently fast relative to the rate at which model performance converges.

Key Challenge: Growing Training Horizons

Recent language-model scaling practices motivate increasing training tokens alongside model size, which often increases gradient updates. The joint scaling limit of width and training horizon is substantially more delicate than the fixed-horizon setting, since the limiting objective itself changes as TT grows.

Model Setup

The model is a linear network with a single trainable hidden matrix:

fn(x;W0,W1,V)=V⊤W1W0xf_n(x; W_0, W_1, V) = V^\top W_1 W_0 x

where W0∈Rn×dW_0 \in \mathbb{R}^{n \times d}, W1∈Rn×nW_1 \in \mathbb{R}^{n \times n}, and V∈RnV \in \mathbb{R}^n. Initialization follows μP scalings: W0,ij∼N(0,d−1)W_{0,ij} \sim \mathcal{N}(0, d^{-1}), W1,ij(0)∼N(0,n−1)W^{(0)}_{1,ij} \sim \mathcal{N}(0, n^{-1}), and Vi∼N(0,n−2)V_i \sim \mathcal{N}(0, n^{-2}).

Key definitions: K=d−1XX⊤∈Rm×mK = d^{-1}XX^\top \in \mathbb{R}^{m \times m} (data Gram matrix), An=∥V∥2HnHn⊤A_n = \|V\|^2 H_n H_n^\top with Hn=XW0⊤H_n = XW_0^\top.

Methodology

Exact Loss Dynamics

The key reduction (Proposition 2.1) shows that training with full-batch gradient descent is equivalent to a scalar problem:

ϕn,T(η):=Ln(W1(T))=(2m)−1(χn(0))⊤Bn(η)2Tχn(0)\phi_{n,T}(\eta) := L_n(W_1^{(T)}) = (2m)^{-1}(\chi_n^{(0)})^\top B_n(\eta)^{2T} \chi_n^{(0)}

where BA(η)=Im−m−1ηAB_A(\eta) = I_m - m^{-1}\eta A and χn(t)=fn(t)(X)−y\chi_n^{(t)} = f_n^{(t)}(X) - y is the residual.

Infinite-Width Reference

The infinite-width objective decomposes spectrally:

ϕ∞,T(η)=c02m+12m∑j=1rcj(1−ηλjm)2T\phi_{\infty,T}(\eta) = \frac{c_0}{2m} + \frac{1}{2m}\sum_{j=1}^{r} c_j \left(1 - \frac{\eta\lambda_j}{m}\right)^{2T}

where λj\lambda_j are distinct positive eigenvalues of KK, cj=∥Pjy∥2c_j = \|P_j y\|^2, and PjP_j are spectral projectors.

Transfer Quantities (Definition 2.2)

  • Loss gap: an,T=∣ϕn,T(ηn,T)−ϕ∞,T(η∞,T)∣a_{n,T} = |\phi_{n,T}(\eta_{n,T}) - \phi_{\infty,T}(\eta_{\infty,T})|
  • Hyperparameter gap: bn,T=∣ηn,T−η∞,T∣b_{n,T} = |\eta_{n,T} - \eta_{\infty,T}|
  • Transfer suboptimality gap: cn,T=ϕ∞,T(ηn,T)−ϕ∞,T(η∞,T)c_{n,T} = \phi_{\infty,T}(\eta_{n,T}) - \phi_{\infty,T}(\eta_{\infty,T})

Fast transfer holds when cn,T=op(an,T)c_{n,T} = o_p(a_{n,T}).

Key Assumptions

Assumption 3.2 (Endpoint activity): The largest and smallest positive eigenvalues λ1\lambda_1 and λr\lambda_r of KK are simple and satisfy ⟨y,u1⟩≠0\langle y, u_1 \rangle \neq 0 and ⟨y,ur⟩≠0\langle y, u_r \rangle \neq 0.

Empirical Validation / Results

Theorem 3.1: Long-Horizon Optimal Learning Rate

As T→∞T \to \infty, the infinite-width optimal learning rate converges to:

η∞,T→η∗y:=2mλ+y+λ−y\eta_{\infty,T} \to \eta_*^y := \frac{2m}{\lambda_+^y + \lambda_-^y}

where λ+y=max⁡j∈Jyλj\lambda_+^y = \max_{j \in J_y} \lambda_j and λ−y=min⁡j∈Jyλj\lambda_-^y = \min_{j \in J_y} \lambda_j. This is related to the classical optimal step size for Richardson iteration.

Lemma 3.3: Joint Initialization Central Limit Theorem

n(χn(0)+y,An−K)→d(Z,Ξ)\sqrt{n}(\chi_n^{(0)} + y, A_n - K) \xrightarrow{d} (Z, \Xi)

where Ξ=αK+XGX⊤\Xi = \alpha K + XGX^\top with α∼N(0,2)\alpha \sim \mathcal{N}(0, 2), GG a symmetric Gaussian matrix, and Z∼N(0,K)Z \sim \mathcal{N}(0, K).

Theorem 3.4: Growing-Horizon Fluctuations and Transfer Rates

Under T→∞T \to \infty and T/n→0T/\sqrt{n} \to 0, with Assumption 3.2:

(ϵnρG,T/κT)−1(ηn,T−η∞,T)→dΩη,∗:=−2m(λ1+λr)2(u1⊤Ξu1+ur⊤Ξur)(\epsilon_n \rho_{G,T}/\kappa_T)^{-1}(\eta_{n,T} - \eta_{\infty,T}) \xrightarrow{d} \Omega_{\eta,*} := -\frac{2m}{(\lambda_1 + \lambda_r)^2}(u_1^\top \Xi u_1 + u_r^\top \Xi u_r) (ϵnρL,T)−1(ϕn,T(ηn,T)−ϕ∞,T(η∞,T))→dΩL,∗:=⟨QL,∗,Ξ⟩F(\epsilon_n \rho_{L,T})^{-1}(\phi_{n,T}(\eta_{n,T}) - \phi_{\infty,T}(\eta_{\infty,T})) \xrightarrow{d} \Omega_{L,*} := \langle Q_{L,*}, \Xi \rangle_F

The resulting rates are:

QuantityRate
an,Ta_{n,T}Θp(Tq∗2T−1n)\Theta_p\left(\frac{T q_*^{2T-1}}{\sqrt{n}}\right)
bn,Tb_{n,T}Θp(1n)\Theta_p\left(\frac{1}{\sqrt{n}}\right)
cn,Tc_{n,T}Θp(T2q∗2T−2n)\Theta_p\left(\frac{T^2 q_*^{2T-2}}{n}\right)

Corollary 3.5: Fast Transfer at Growing Horizons

cn,Tan,T=Θp(Tq∗n)→p0\frac{c_{n,T}}{a_{n,T}} = \Theta_p\left(\frac{T}{q_* \sqrt{n}}\right) \xrightarrow{p} 0

when T=o(n)T = o(\sqrt{n}).

Theorem 3.6: Fixed-Horizon Rates

For fixed TT, under nondegeneracy (∣Jy∣≥2|J_y| \geq 2): an,T=Θp(n−1/2)a_{n,T} = \Theta_p(n^{-1/2}), bn,T=Θp(n−1/2)b_{n,T} = \Theta_p(n^{-1/2}), cn,T=Θp(n−1)c_{n,T} = \Theta_p(n^{-1}).

Proposition 3.8: Counterexample to Fast Transfer

With a single positive eigenvalue (XX⊤=ImXX^\top = I_m):

an,T=Θp(n−T),bn,T=Θp(n−1/2),cn,T=Θp(n−T),cn,T/an,T=Θp(1)a_{n,T} = \Theta_p(n^{-T}), \quad b_{n,T} = \Theta_p(n^{-1/2}), \quad c_{n,T} = \Theta_p(n^{-T}), \quad c_{n,T}/a_{n,T} = \Theta_p(1)

Fast transfer fails because the loss gap and transfer penalty scale identically.

Theoretical and Practical Implications

Local Perturbation Formulation (Section 4.2)

The paper abstracts the proof structure into a general framework with three scales:

  • ρL,T\rho_{L,T}: Loss-value response to finite-width perturbations
  • ρG,T\rho_{G,T}: Loss-gradient response to finite-width perturbations
  • κT\kappa_T: Local curvature at the infinite-width optimizer

The transfer penalty satisfies:

cn,T=Op(ϵn2ρG,T2κT)c_{n,T} = O_p\left(\frac{\epsilon_n^2 \rho_{G,T}^2}{\kappa_T}\right)

Fast transfer follows when ϵnρG,T2/(κTρL,T)→0\epsilon_n \rho_{G,T}^2 / (\kappa_T \rho_{L,T}) \to 0.

Key Insights

  1. Spectral dependence: Different spectra lead to different transfer rates—contrasting with Wen et al. (2026), where rates are more uniform.

  2. Directional misalignment: The loss-value and loss-gradient responses to the same perturbation need not be aligned. Some perturbation directions change the optimized loss at leading order while producing no leading-order shift in the optimal learning rate.

  3. Stability vs. optimality: The largest eigenvalue determines the stability boundary (2m/λ12m/\lambda_1), but the long-horizon optimum depends on both endpoints: η∗=2m/(λ1+λr)\eta_* = 2m/(\lambda_1 + \lambda_r).

  4. Low-rank concentration is insufficient: The top-1 structure of the gradient update holds in both fast-transfer and failure regimes, showing that structural properties alone don't determine transfer quality.

Conclusion

Main takeaways:

  • Fast learning-rate transfer persists at growing horizons when T=o(n)T = o(\sqrt{n}) under endpoint activity assumptions
  • Transfer rates decompose cleanly into width-scale (ϵn\epsilon_n) and horizon-dependent sensitivity scales (ρL,T\rho_{L,T}, ρG,T\rho_{G,T}, κT\kappa_T)
  • Nondegenerate first-order loss sensitivity is essential for fast transfer

Limitations and future directions:

  • Restricted to linear networks with a single trainable layer
  • Extension to nonlinear/deep networks is an important next step
  • Current framework relies on nondegenerate first-order sensitivities; developing higher-order analysis would help understand failure cases like Proposition 3.8
  • Studying transfer across multiple scaling axes (width and depth) simultaneously remains open

Related papers