# Normalize-Then-Precondition: A Hierarchical Approach to Marginal Scale and Interaction Geometry for LLM Training

> NormPre separates LLM update normalization from spectral preconditioning, consistently beating AdamW, Muon, and MANO in pretraining while cutting optimizer latency by up to 67%.

- **Source:** [arXiv](https://arxiv.org/abs/2609.36692)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/Nm8Co5
- **Whiteboard:** https://picx.dev/p/Nm8Co5/image

## Summary

## Summary (Overview)

- **Core Contribution**: Introduces the **Normalize-Then-Precondition** framework, a hierarchical approach for LLM training that first normalizes updates using marginal-scale information (diagonal Gram) and then applies spectral preconditioning to refine directional interaction geometry.
- **Two Novel Optimizers**: Proposes **NormPre-G** (global spectral preconditioning via Newton-Schulz iterations) and **NormPre-L** (localized spectral preconditioning via randomized sketching), both built on the framework with alternating row/column normalization and consistent update RMS scaling.
- **Theoretical Guarantees**: Establishes $O(T^{-1/2})$ convergence guarantees for simplified versions of NormPre, extending to sketch-based variants under controlled approximation error.
- **Empirical Superiority**: Both variants consistently outperform AdamW, Muon, and MANO across GPT-2 Small, LLaMA (up to 1.3B), and Qwen3 (up to 1.7B) pretraining under matched training budgets, with NormPre-G achieving the lowest validation loss.
- **Performance-Efficiency Trade-off**: NormPre-L reduces optimizer latency by up to 67% compared to Muon while maintaining strong optimization gains, offering practitioners flexible choices.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Large language model (LLM) training demands optimizers that balance performance, efficiency, and scalability. Recent advances exploit parameter structure at multiple granularities:
- **Coordinate-wise**: AdamW, Lion, Sophia
- **Block-wise**: Adam-mini, Blockwise LR
- **Matrix-level**: Muon (orthogonalization), MANO (normalization), Shampoo, SOAP

### Key Theoretical Insight: Gram Representations

The paper reveals a shared structure between orthogonalization and normalization through Gram matrix representations. For a matrix update $X$, the transformations take the forms:

$$\Phi = (XX^\top)^{-1/2}X \quad \text{(Full-Gram Representation)}$$

$$\Psi = \left[\text{Diag}(\text{diag}(XX^\top))\right]^{-1/2}X \quad \text{(Diagonal-Gram Representation)}$$

where $\text{Diag}(\text{diag}(\cdot))$ retains only diagonal entries.

**Critical observation**: The full Gram matrix jointly encodes marginal scales (diagonal entries $\|x_i\|_2^2$) and cross-row interactions (off-diagonal entries $\langle x_i, x_j \rangle$). As scale gaps grow extreme (e.g., one row norm of $10^4$ vs. $10^{-3}$), the leading Gram eigenvector becomes dominated by marginal scales, obscuring directional interactions. This motivates separating scale normalization from interaction processing.

### Theoretical Foundation: Spectral Steepest Descent

**Global preconditioning** solves the spectral steepest descent problem:

$$T_G \in \arg\max_{\|T\|_{op} \leq 1} \langle \Psi, T \rangle_F$$

with solution $T_G = P_G \Psi = \text{msign}(\Psi)$, where $P_G = \Gamma(X)^{-1/2} = \tilde{U}\Lambda^{-1/2}\tilde{U}^\top$.

**Localized preconditioning** retains $\Psi$ as reference and solves a regularized problem:

$$T_+ := \arg\min_{\|T\|_{op} \leq 1} \frac{1}{2}\|T - \Psi\|_F^2$$

with solution $T_+ = P_+ \Psi$ where $P_+ = I_m + \tilde{U}_A(\Lambda_A^{-1/2} - I)\tilde{U}_A^\top$ for active set $A := \{i : \lambda_i > 1\}$.

---

## Methodology

### The NormPre Optimizer (Algorithm 1)

The optimizer follows a structured pipeline per step:

1. **Momentum update**: $M_t \leftarrow \mu M_{t-1} + G_t$
2. **Alternating orientation**: $k_t \leftarrow t \bmod 2$ (switches between row/column views)
3. **Relaxed tangent momentum**: $x_{t,i} := \bar{m}_{t,i} - \langle \bar{m}_{t,i}, \bar{w}_{t,i}\rangle \bar{w}_{t,i}$ (separates tangent/radial components)
4. **Diagonal-Gram normalization**: $\Psi_t \leftarrow D(X_t)^{-1}X_t$
5. **Spectral preconditioning** (variant-dependent):
   - **NormPre-G**: $T_t \leftarrow \text{msign}(\Psi_t)$ via 5 Newton-Schulz iterations
   - **NormPre-L**: Extract top-$r$ active eigenspace, then $T_t \leftarrow (I + \tilde{U}_{C_t}(\Lambda_{C_t}^{-1/2} - I)\tilde{U}_{C_t}^\top)\Psi_t$
6. **Consistent update RMS**: $W_{t+1} \leftarrow W_t - \eta_t(R_{k_t}(T_t) + \lambda_{wd}W_t)$ with target RMS of 0.2

### Scalable Implementations

**NormPre-G** uses Newton-Schulz iterations (same complexity as Muon): $O(mn + qmns)$ where $s = \min\{m,n\}$.

**NormPre-L** offers two implementations:
- **Exact**: Full eigendecomposition of $\Gamma_t = \Psi_t\Psi_t^\top$: $O(m^2n + m^3)$
- **Sketch-based**: Randomized sketching with Rayleigh-Ritz extraction: $O((p+1)mn\ell + (m+n)\ell^2 + \ell^3)$ where $\ell = \min\{m, r+o\}$

### Convergence Guarantee

**Theorem 1** (Convergence without momentum): Under $L$-smoothness, bounded radial ratios, and spectral factor conditions, choosing $\eta = C/\sqrt{T+1}$ yields:

$$\min_{0 \leq t \leq T} \|\nabla L(W_t)\|_F \leq \frac{1}{\sqrt{T+1}} \left(\frac{\Delta_0 \max\{m,n\}^{1/2}}{\epsilon\gamma C} + \frac{LC \max\{m,n\}^{3/2}}{2\epsilon\gamma}\right)$$

The optimal rate is $\min_{0 \leq t \leq T} \|\nabla L(W_t)\|_F \leq \frac{\max\{m,n\}\sqrt{2L\Delta_0}}{\epsilon\gamma\sqrt{T+1}}$.

---

## Empirical Validation / Results

### Scaling Experiments

**Table 1: Validation loss across model scales, architectures, and datasets** (lower is better):

| Setting | | | | **NormPre-G** | **NormPre-L** |
|---------|---------|---------|---------|---------|---------|
| **Model** | **Dataset** | **AdamW** | **Muon** | **MANO** | | |
| GPT-2 Small | OpenWebText | 3.1444 ± 0.0054 | 3.1064 ± 0.0054 | 3.1156 ± 0.0057 | **3.0667** ± 0.0049 | 3.0826 ± 0.0034 |
| LLaMA-130M | C4 | 3.1363 | 3.1019 | 3.1102 | **3.0737** | 3.0930 |
| LLaMA-350M | C4 | 3.0378 | 2.9999 | 2.9946 | **2.9678** | 2.9758 |
| LLaMA-1.3B | C4 | 2.9385 | 2.9037 | 2.8963 | **2.8571** | 2.8662 |
| Qwen3-0.6B | Pile | 2.8956 | 2.8335 | 2.8382 | **2.7967** | 2.8066 |
| Qwen3-1.7B | Pile | 2.6758 | 2.6408 | 2.6205 | **2.5868** | 2.5897 |

**Key results**: Both variants outperform all baselines across every setting. NormPre-G achieves the lowest validation loss universally. vs. Muon (strongest baseline), NormPre-G reduces loss by 0.0397 (GPT-2 Small) to 0.0392 (LLaMA-1.3B).

### Training Efficiency (Table 2)

| Model | Optimizer | Optimizer Latency | E2E Step Time | Throughput | Peak Memory |
|-------|-----------|-------------------|---------------|------------|-------------|
| GPT-2 Small | NormPre-G | 119.3 ms (+5.76% vs Muon) | 3495.0 ms (+0.44%) | 150.01 k tok/s | 14.62 GiB |
| GPT-2 Small | NormPre-L | 69.8 ms (−38.12% vs Muon) | 3439.4 ms (−1.16%) | 152.44 k tok/s | 14.61 GiB |
| LLaMA-1.3B | NormPre-G | 1054.2 ms (+4.37% vs Muon) | 23141.8 ms (+0.08%) | 22.66 k tok/s | 37.83 GiB |
| LLaMA-1.3B | NormPre-L | 333.3 ms (−67.00% vs Muon) | 22399.7 ms (−3.13%) | 23.41 k tok/s | 37.84 GiB |

### Spectral Dynamics

Key findings from spectral analysis:
1. **Normalization yields anisotropic base updates**: Equalizing row norms redistributes spectral mass, raising eigenvalues across multiple ranks—anisotropy persists throughout training, justifying further refinement.
2. **Localized preconditioning captures dominant energy**: Top-32 modes ($\lambda_i > 1$) consistently account for approximately **62%–67%** of global spectral transformation energy across checkpoints.

---

## Theoretical and Practical Implications

### Theoretical Implications

- **Hierarchical geometric organization**: The framework formally separates marginal-scale information (diagonal Gram) from interaction information (off-diagonal Gram), providing a principled alternative to joint geometric transformations.
- **Unified perspective**: Global and localized preconditioning emerge as solutions to distinct optimization problems—spectral steepest descent vs. regularized steepest descent with leading mode selection—unifying previously disparate approaches.
- **Convergence theory**: The $O(T^{-1/2})$ rate matches standard stochastic optimization guarantees while accounting for the geometric structure of matrix updates.

### Practical Implications

- **Flexible performance-efficiency trade-off**: NormPre-L with rank $r = 32$ offers strong gains at substantially lower computational cost (up to 67% latency reduction vs. Muon), while NormPre-G pushes absolute performance limits.
- **Direct drop-in replacement**: Both variants maintain the Muon convention of 0.2 update RMS, enabling shared hyperparameters with AdamW and Muon.
- **Scalability**: Sketch-based eigenspace extraction makes localized preconditioning feasible for large matrices where full eigendecomposition is prohibitive.

---

## Conclusion

### Main Takeaways

This paper formalizes the **Normalize-Then-Precondition** framework, establishing that:
1. Marginal-scale normalization (via diagonal Gram) should precede interaction processing (via spectral preconditioning)
2. Both full-spectrum (global) and targeted (localized) spectral transformations offer distinct advantages
3. The resulting NormPre optimizers consistently outperform established baselines across multiple architectures and scales

### Future Directions

The authors highlight three promising avenues:
- **(a)** Hardware-aware implementations to improve spectral transformation efficiency
- **(b)** Scaling empirical validation beyond 1.7B parameters
- **(c)** Exploring adaptive spectral schemes to narrow the gap between localized and global preconditioning while preserving efficiency

### Key Formula Summary

The central theoretical contributions are:

**Full-Gram (Muon)**: $\Phi = (XX^\top)^{-1/2}X$

**Diagonal-Gram (Normalization)**: $\Psi = \left[\text{Diag}(\text{diag}(XX^\top))\right]^{-1/2}X$

**Global preconditioner**: $P_G = \Gamma(X)^{-1/2} = \tilde{U}\Lambda^{-1/2}\tilde{U}^\top$

**Localized preconditioner**: $P_L = I_m + \tilde{U}_C(\Lambda_C^{-1/2} - I)\tilde{U}_C^\top$

---

_Markdown view of https://picx.dev/p/Nm8Co5, served by PicX — AI-generated visual whiteboard summaries of research papers._
