Full text not available for this paper
Summary (Overview)
- Core Contribution: Introduces the Normalize-Then-Precondition framework, a hierarchical approach for LLM training that first normalizes updates using marginal-scale information (diagonal Gram) and then applies spectral preconditioning to refine directional interaction geometry.
- Two Novel Optimizers: Proposes NormPre-G (global spectral preconditioning via Newton-Schulz iterations) and NormPre-L (localized spectral preconditioning via randomized sketching), both built on the framework with alternating row/column normalization and consistent update RMS scaling.
- Theoretical Guarantees: Establishes convergence guarantees for simplified versions of NormPre, extending to sketch-based variants under controlled approximation error.
- Empirical Superiority: Both variants consistently outperform AdamW, Muon, and MANO across GPT-2 Small, LLaMA (up to 1.3B), and Qwen3 (up to 1.7B) pretraining under matched training budgets, with NormPre-G achieving the lowest validation loss.
- Performance-Efficiency Trade-off: NormPre-L reduces optimizer latency by up to 67% compared to Muon while maintaining strong optimization gains, offering practitioners flexible choices.
Introduction and Theoretical Foundation
Background and Motivation
Large language model (LLM) training demands optimizers that balance performance, efficiency, and scalability. Recent advances exploit parameter structure at multiple granularities:
- Coordinate-wise: AdamW, Lion, Sophia
- Block-wise: Adam-mini, Blockwise LR
- Matrix-level: Muon (orthogonalization), MANO (normalization), Shampoo, SOAP
Key Theoretical Insight: Gram Representations
The paper reveals a shared structure between orthogonalization and normalization through Gram matrix representations. For a matrix update , the transformations take the forms:
where retains only diagonal entries.
Critical observation: The full Gram matrix jointly encodes marginal scales (diagonal entries ) and cross-row interactions (off-diagonal entries ). As scale gaps grow extreme (e.g., one row norm of vs. ), the leading Gram eigenvector becomes dominated by marginal scales, obscuring directional interactions. This motivates separating scale normalization from interaction processing.
Theoretical Foundation: Spectral Steepest Descent
Global preconditioning solves the spectral steepest descent problem:
with solution , where .
Localized preconditioning retains as reference and solves a regularized problem:
with solution where for active set .
Methodology
The NormPre Optimizer (Algorithm 1)
The optimizer follows a structured pipeline per step:
- Momentum update:
- Alternating orientation: (switches between row/column views)
- Relaxed tangent momentum: (separates tangent/radial components)
- Diagonal-Gram normalization:
- Spectral preconditioning (variant-dependent):
- NormPre-G: via 5 Newton-Schulz iterations
- NormPre-L: Extract top- active eigenspace, then
- Consistent update RMS: with target RMS of 0.2
Scalable Implementations
NormPre-G uses Newton-Schulz iterations (same complexity as Muon): where .
NormPre-L offers two implementations:
- Exact: Full eigendecomposition of :
- Sketch-based: Randomized sketching with Rayleigh-Ritz extraction: where
Convergence Guarantee
Theorem 1 (Convergence without momentum): Under -smoothness, bounded radial ratios, and spectral factor conditions, choosing yields:
The optimal rate is .
Empirical Validation / Results
Scaling Experiments
Table 1: Validation loss across model scales, architectures, and datasets (lower is better):
| Setting | NormPre-G | NormPre-L | |||
|---|---|---|---|---|---|
| Model | Dataset | AdamW | Muon | MANO | |
| GPT-2 Small | OpenWebText | 3.1444 ± 0.0054 | 3.1064 ± 0.0054 | 3.1156 ± 0.0057 | 3.0667 ± 0.0049 |
| LLaMA-130M | C4 | 3.1363 | 3.1019 | 3.1102 | 3.0737 |
| LLaMA-350M | C4 | 3.0378 | 2.9999 | 2.9946 | 2.9678 |
| LLaMA-1.3B | C4 | 2.9385 | 2.9037 | 2.8963 | 2.8571 |
| Qwen3-0.6B | Pile | 2.8956 | 2.8335 | 2.8382 | 2.7967 |
| Qwen3-1.7B | Pile | 2.6758 | 2.6408 | 2.6205 | 2.5868 |
Key results: Both variants outperform all baselines across every setting. NormPre-G achieves the lowest validation loss universally. vs. Muon (strongest baseline), NormPre-G reduces loss by 0.0397 (GPT-2 Small) to 0.0392 (LLaMA-1.3B).
Training Efficiency (Table 2)
| Model | Optimizer | Optimizer Latency | E2E Step Time | Throughput | Peak Memory |
|---|---|---|---|---|---|
| GPT-2 Small | NormPre-G | 119.3 ms (+5.76% vs Muon) | 3495.0 ms (+0.44%) | 150.01 k tok/s | 14.62 GiB |
| GPT-2 Small | NormPre-L | 69.8 ms (−38.12% vs Muon) | 3439.4 ms (−1.16%) | 152.44 k tok/s | 14.61 GiB |
| LLaMA-1.3B | NormPre-G | 1054.2 ms (+4.37% vs Muon) | 23141.8 ms (+0.08%) | 22.66 k tok/s | 37.83 GiB |
| LLaMA-1.3B | NormPre-L | 333.3 ms (−67.00% vs Muon) | 22399.7 ms (−3.13%) | 23.41 k tok/s | 37.84 GiB |
Spectral Dynamics
Key findings from spectral analysis:
- Normalization yields anisotropic base updates: Equalizing row norms redistributes spectral mass, raising eigenvalues across multiple ranks—anisotropy persists throughout training, justifying further refinement.
- Localized preconditioning captures dominant energy: Top-32 modes () consistently account for approximately 62%–67% of global spectral transformation energy across checkpoints.
Theoretical and Practical Implications
Theoretical Implications
- Hierarchical geometric organization: The framework formally separates marginal-scale information (diagonal Gram) from interaction information (off-diagonal Gram), providing a principled alternative to joint geometric transformations.
- Unified perspective: Global and localized preconditioning emerge as solutions to distinct optimization problems—spectral steepest descent vs. regularized steepest descent with leading mode selection—unifying previously disparate approaches.
- Convergence theory: The rate matches standard stochastic optimization guarantees while accounting for the geometric structure of matrix updates.
Practical Implications
- Flexible performance-efficiency trade-off: NormPre-L with rank offers strong gains at substantially lower computational cost (up to 67% latency reduction vs. Muon), while NormPre-G pushes absolute performance limits.
- Direct drop-in replacement: Both variants maintain the Muon convention of 0.2 update RMS, enabling shared hyperparameters with AdamW and Muon.
- Scalability: Sketch-based eigenspace extraction makes localized preconditioning feasible for large matrices where full eigendecomposition is prohibitive.
Conclusion
Main Takeaways
This paper formalizes the Normalize-Then-Precondition framework, establishing that:
- Marginal-scale normalization (via diagonal Gram) should precede interaction processing (via spectral preconditioning)
- Both full-spectrum (global) and targeted (localized) spectral transformations offer distinct advantages
- The resulting NormPre optimizers consistently outperform established baselines across multiple architectures and scales
Future Directions
The authors highlight three promising avenues:
- (a) Hardware-aware implementations to improve spectral transformation efficiency
- (b) Scaling empirical validation beyond 1.7B parameters
- (c) Exploring adaptive spectral schemes to narrow the gap between localized and global preconditioning while preserving efficiency
Key Formula Summary
The central theoretical contributions are:
Full-Gram (Muon):
Diagonal-Gram (Normalization):
Global preconditioner:
Localized preconditioner:
Related papers
- Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons
Optimal learning-rate warmup duration scales with training horizon only at high peak learning rates, following a regime-dependent law predictable from three short runs.
- Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion
Local attention-output reconstruction gains do not guarantee final-model fidelity, as residual completion can improve local error while worsening dense-model KL divergence.
- How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Under matched conditions, a minimal-harness coding agent matches or outperforms state-of-the-art MLE harnesses, with performance driven by the LLM backbone and execution environment, not scaffolding.