# Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons

> Fast learning-rate transfer holds at growing training horizons when T grows slower than sqrt(n), but requires nondegenerate first-order loss sensitivity to avoid spectral-dependent failures.

- **Source:** [arXiv](https://arxiv.org/abs/2609.35029)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/wUwhUw
- **Whiteboard:** https://picx.dev/p/wUwhUw/image

## Summary

## Summary (Overview)

- **Core contribution**: This paper proves fast learning-rate transfer in shallow linear networks when the training horizon grows jointly with width, specifically when $T = o(\sqrt{n})$, under spectral assumptions on the data Gram matrix.

- **Theoretical framework**: The authors leverage the notion of "fast hyperparameter transfer" (Ghosh et al., 2026) and show that convergence of the optimal learning rate alone is insufficient; what matters is the rate of convergence relative to the loss convergence rate.

- **Key rate characterization**: The transfer rates are decomposed into three horizon-dependent quantities—loss-value sensitivity ($\rho_{L,T}$), loss-gradient sensitivity ($\rho_{G,T}$), and local curvature ($\kappa_T$)—combined with the finite-width perturbation scale $\epsilon_n = n^{-1/2}$.

- **Distributional results**: The paper derives limiting distributions for both the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix.

- **Counterexample**: A fixed-horizon counterexample (single positive eigenvalue) shows fast transfer can fail even when the hyperparameter gap converges at the standard $n^{-1/2}$ rate, demonstrating that nondegenerate first-order loss sensitivity is essential.

## Introduction and Theoretical Foundation

### Background

The paper addresses the growing cost of hyperparameter tuning for large neural networks. **Hyperparameter transfer** reduces this cost by reusing hyperparameters tuned on small proxy models when training larger models. The framework builds on:

- **μP (Maximal Update Parameterization)** (Yang and Hu, 2021): Specifies width-dependent initialization and learning-rate scalings that preserve nontrivial feature learning in the infinite-width limit.
- **μTransfer** (Yang et al., 2021): Tune hyperparameters on a small proxy and transfer to larger models without retuning.
- **Fast hyperparameter transfer** (Ghosh et al., 2026): Formalizes when transfer is effective—what matters is whether the optimal hyperparameter converges sufficiently fast *relative to* the rate at which model performance converges.

### Key Challenge: Growing Training Horizons

Recent language-model scaling practices motivate increasing training tokens alongside model size, which often increases gradient updates. The joint scaling limit of width and training horizon is substantially more delicate than the fixed-horizon setting, since the limiting objective itself changes as $T$ grows.

### Model Setup

The model is a linear network with a single trainable hidden matrix:

$$f_n(x; W_0, W_1, V) = V^\top W_1 W_0 x$$

where $W_0 \in \mathbb{R}^{n \times d}$, $W_1 \in \mathbb{R}^{n \times n}$, and $V \in \mathbb{R}^n$. Initialization follows μP scalings: $W_{0,ij} \sim \mathcal{N}(0, d^{-1})$, $W^{(0)}_{1,ij} \sim \mathcal{N}(0, n^{-1})$, and $V_i \sim \mathcal{N}(0, n^{-2})$.

**Key definitions**: $K = d^{-1}XX^\top \in \mathbb{R}^{m \times m}$ (data Gram matrix), $A_n = \|V\|^2 H_n H_n^\top$ with $H_n = XW_0^\top$.

## Methodology

### Exact Loss Dynamics

The key reduction (Proposition 2.1) shows that training with full-batch gradient descent is equivalent to a scalar problem:

$$\phi_{n,T}(\eta) := L_n(W_1^{(T)}) = (2m)^{-1}(\chi_n^{(0)})^\top B_n(\eta)^{2T} \chi_n^{(0)}$$

where $B_A(\eta) = I_m - m^{-1}\eta A$ and $\chi_n^{(t)} = f_n^{(t)}(X) - y$ is the residual.

### Infinite-Width Reference

The infinite-width objective decomposes spectrally:

$$\phi_{\infty,T}(\eta) = \frac{c_0}{2m} + \frac{1}{2m}\sum_{j=1}^{r} c_j \left(1 - \frac{\eta\lambda_j}{m}\right)^{2T}$$

where $\lambda_j$ are distinct positive eigenvalues of $K$, $c_j = \|P_j y\|^2$, and $P_j$ are spectral projectors.

### Transfer Quantities (Definition 2.2)

- **Loss gap**: $a_{n,T} = |\phi_{n,T}(\eta_{n,T}) - \phi_{\infty,T}(\eta_{\infty,T})|$
- **Hyperparameter gap**: $b_{n,T} = |\eta_{n,T} - \eta_{\infty,T}|$
- **Transfer suboptimality gap**: $c_{n,T} = \phi_{\infty,T}(\eta_{n,T}) - \phi_{\infty,T}(\eta_{\infty,T})$

**Fast transfer** holds when $c_{n,T} = o_p(a_{n,T})$.

### Key Assumptions

**Assumption 3.2 (Endpoint activity)**: The largest and smallest positive eigenvalues $\lambda_1$ and $\lambda_r$ of $K$ are simple and satisfy $\langle y, u_1 \rangle \neq 0$ and $\langle y, u_r \rangle \neq 0$.

## Empirical Validation / Results

### Theorem 3.1: Long-Horizon Optimal Learning Rate

As $T \to \infty$, the infinite-width optimal learning rate converges to:

$$\eta_{\infty,T} \to \eta_*^y := \frac{2m}{\lambda_+^y + \lambda_-^y}$$

where $\lambda_+^y = \max_{j \in J_y} \lambda_j$ and $\lambda_-^y = \min_{j \in J_y} \lambda_j$. This is related to the classical optimal step size for Richardson iteration.

### Lemma 3.3: Joint Initialization Central Limit Theorem

$$\sqrt{n}(\chi_n^{(0)} + y, A_n - K) \xrightarrow{d} (Z, \Xi)$$

where $\Xi = \alpha K + XGX^\top$ with $\alpha \sim \mathcal{N}(0, 2)$, $G$ a symmetric Gaussian matrix, and $Z \sim \mathcal{N}(0, K)$.

### Theorem 3.4: Growing-Horizon Fluctuations and Transfer Rates

Under $T \to \infty$ and $T/\sqrt{n} \to 0$, with Assumption 3.2:

$$(\epsilon_n \rho_{G,T}/\kappa_T)^{-1}(\eta_{n,T} - \eta_{\infty,T}) \xrightarrow{d} \Omega_{\eta,*} := -\frac{2m}{(\lambda_1 + \lambda_r)^2}(u_1^\top \Xi u_1 + u_r^\top \Xi u_r)$$

$$(\epsilon_n \rho_{L,T})^{-1}(\phi_{n,T}(\eta_{n,T}) - \phi_{\infty,T}(\eta_{\infty,T})) \xrightarrow{d} \Omega_{L,*} := \langle Q_{L,*}, \Xi \rangle_F$$

The resulting rates are:

| Quantity | Rate |
|----------|------|
| $a_{n,T}$ | $\Theta_p\left(\frac{T q_*^{2T-1}}{\sqrt{n}}\right)$ |
| $b_{n,T}$ | $\Theta_p\left(\frac{1}{\sqrt{n}}\right)$ |
| $c_{n,T}$ | $\Theta_p\left(\frac{T^2 q_*^{2T-2}}{n}\right)$ |

### Corollary 3.5: Fast Transfer at Growing Horizons

$$\frac{c_{n,T}}{a_{n,T}} = \Theta_p\left(\frac{T}{q_* \sqrt{n}}\right) \xrightarrow{p} 0$$

when $T = o(\sqrt{n})$.

### Theorem 3.6: Fixed-Horizon Rates

For fixed $T$, under nondegeneracy ($|J_y| \geq 2$): $a_{n,T} = \Theta_p(n^{-1/2})$, $b_{n,T} = \Theta_p(n^{-1/2})$, $c_{n,T} = \Theta_p(n^{-1})$.

### Proposition 3.8: Counterexample to Fast Transfer

With a single positive eigenvalue ($XX^\top = I_m$):

$$a_{n,T} = \Theta_p(n^{-T}), \quad b_{n,T} = \Theta_p(n^{-1/2}), \quad c_{n,T} = \Theta_p(n^{-T}), \quad c_{n,T}/a_{n,T} = \Theta_p(1)$$

Fast transfer **fails** because the loss gap and transfer penalty scale identically.

## Theoretical and Practical Implications

### Local Perturbation Formulation (Section 4.2)

The paper abstracts the proof structure into a general framework with three scales:

- **$\rho_{L,T}$**: Loss-value response to finite-width perturbations
- **$\rho_{G,T}$**: Loss-gradient response to finite-width perturbations  
- **$\kappa_T$**: Local curvature at the infinite-width optimizer

The transfer penalty satisfies:

$$c_{n,T} = O_p\left(\frac{\epsilon_n^2 \rho_{G,T}^2}{\kappa_T}\right)$$

Fast transfer follows when $\epsilon_n \rho_{G,T}^2 / (\kappa_T \rho_{L,T}) \to 0$.

### Key Insights

1. **Spectral dependence**: Different spectra lead to different transfer rates—contrasting with Wen et al. (2026), where rates are more uniform.

2. **Directional misalignment**: The loss-value and loss-gradient responses to the same perturbation need not be aligned. Some perturbation directions change the optimized loss at leading order while producing no leading-order shift in the optimal learning rate.

3. **Stability vs. optimality**: The largest eigenvalue determines the stability boundary ($2m/\lambda_1$), but the long-horizon optimum depends on both endpoints: $\eta_* = 2m/(\lambda_1 + \lambda_r)$.

4. **Low-rank concentration is insufficient**: The top-1 structure of the gradient update holds in both fast-transfer and failure regimes, showing that structural properties alone don't determine transfer quality.

## Conclusion

**Main takeaways**:
- Fast learning-rate transfer persists at growing horizons when $T = o(\sqrt{n})$ under endpoint activity assumptions
- Transfer rates decompose cleanly into width-scale ($\epsilon_n$) and horizon-dependent sensitivity scales ($\rho_{L,T}$, $\rho_{G,T}$, $\kappa_T$)
- Nondegenerate first-order loss sensitivity is essential for fast transfer

**Limitations and future directions**:
- Restricted to linear networks with a single trainable layer
- Extension to nonlinear/deep networks is an important next step
- Current framework relies on nondegenerate first-order sensitivities; developing higher-order analysis would help understand failure cases like Proposition 3.8
- Studying transfer across multiple scaling axes (width and depth) simultaneously remains open

---

_Markdown view of https://picx.dev/p/wUwhUw, served by PicX — AI-generated visual whiteboard summaries of research papers._
