Scaling Laws & Training MethodsIssue 9Oct 3 – 10, 2026

Batch size becomes a new axis for optimizer evaluation

Highlights of this issue

The most prominent theme of this issue is establishing "batch size" as an axis for optimizer evaluation as important as training duration. The Best Optimizer Depends on Batch Size (https://arxiv.org/abs/2610.08975) systematically proves that even when independently tuning for each batch size, optimizer rankings still invert with batch—SOAP is optimal for small batches, Shampoo for large batches, and Muon has no principled scaling law that transfers across settings. The authors use a noisy quadratic model to show that the bias-variance tradeoff of preconditioning can cause ranking inversions, and provide direction-dependent scaling laws: the higher the curvature-to-noise ratio (CNR), the scaling exponent shifts from square root to linear. This extends the work from issue 7 that used training duration as an optimizer evaluation axis, to the batch size axis, which was previously fixed by default or scaled heuristically, directly challenging the common practice of evaluating optimizers under a single batch size.

On the hyperparameter scaling law side, MoE sparsity is incorporated as a predictive dimension. Hyperparameter Scaling Laws Across MoE Sparsity (https://arxiv.org/abs/2609.08690) is the first to include the activation rate A as an independent multiplicative power-law factor in the unified form h*(X,A)=k X^γ A^δ, validated with 1800 runs in the 1/64 ultra-sparse regime, and shows that the optimal learning rate scales with compute C, batch scales with token count D, and A multiplicatively modifies both. This directly responds to the active line of research on muP-for-MoE and the reliability of hyperparameter scaling law extrapolation. From Spectra to Joint Schedules in LLM Pre-training (https://arxiv.org/abs/2609.40148) provides, from a spectral mechanism perspective, criteria for the impact of joint learning rate-batch size schedules on loss power laws: whether component power laws hold is determined by cumulative weighted spectral quality rather than per-coordinate power laws, and characterizes the boundaries where schedules preserve, alter, or break clean power laws, validated on 300M nanoGPT with B/η path equivalence and cross-schedule zero-refit predictions.

Spectral shaping and stability diagnostics also see progress. Does Muon Need Fine-Grained Spectral Shaping (https://arxiv.org/abs/2610.07497) uses spectral diagnostics to find that 94–97% of singular modes lie below the noise edge but are collectively aligned, and proposes BulkBoost, which uses only two bands (bulk/spike) with a single shared gain, proving that coarse-grained reweighting can match fine-grained power-law spectral shaping. Muon Sublates the Edge of Stability (https://arxiv.org/abs/2609.34915) decomposes Muon's edge of stability into two independent signals: a conditionally loss-neutral boundary and temporal direction alignment, proving that Muon breaks the classical picture where the two are coupled in GD. σTransfer (https://arxiv.org/abs/2610.11668) generalizes muP's hyperparameter transfer from learning rate to the prior precision of Laplace approximation, selecting precision on small networks and zero-shot transferring to large networks, and supports decision transfer for active learning, OOD detection, and more.

Community and updates

  • Scaling Laws, Honestly (https://www.completeskeptic.com/p/scaling-laws-honestly) revisits, in an opinion piece, the methodological flaws of the original Kaplan scaling laws: fixed token budgets and learning rate schedules cause small models to be relatively over-trained, leading to a recipe of "too large models, too little data" being used for years. Complementing the debate in issue 8 on The Death of Chinchilla regarding the applicability of compute-optimal ratios, this piece focuses on historical methodological flaws rather than emergent abilities, contributing to the ongoing debate on the reliability of scaling law extrapolation.

Open questions

  1. The Best Optimizer Depends on Batch Size proves that optimizer rankings under a single batch size are unreliable, but the overtraining axis evaluation in issue 7's Optimizer Memory Schedules is also only conducted under fixed batch sizes. Can the batch size and training duration axes be unified into a joint optimizer evaluation protocol, or will their interaction further invert rankings?
  2. BulkBoost's two-band reweighting matches fine-grained spectral shaping in the 410M continued training setting, but issue 4's Spectral Allocation per-rank measurements show a stable spectrum with "head anchoring, body amplification." Does coarse-grained two-band still hold for from-scratch pretraining and larger scales, or does the optimal granularity of spectral shaping itself vary with scale?
  3. Hyperparameter Scaling Laws Across MoE Sparsity incorporates the activation rate A multiplicatively into learning rate and batch laws, but muP-for-MoE's width transfer (issue 1's MSSP) and token dimension scaling (issue 3) do not explicitly handle A. Can the activation rate be incorporated into the muP framework as an independent coordinate, or is it just a function approximation of total/active parameter counts?
  4. σTransfer generalizes muP transfer from learning rate to Laplace prior precision, but the theoretical guarantees are limited to fixed-depth ReLU MLPs. Is hyperparameter transfer for uncertainty estimation equally stable in Transformers and at larger scales, or does it share the same failure boundary as learning rate transfer?

Papers in this issue

  1. Optimizer rankings reverse across batch sizes even after extensive tuning, driven by a bias-variance tradeoff in preconditioning that makes single-exponent scaling rules fundamentally limited.

    Editor's note

    First systematic proof that batch size inverts optimizer rankings (SOAP optimal for small batches, Shampoo for large batches), and provides direction-dependent batch scaling laws: the higher the CNR, the scaling exponent shifts from square root to linear. Compared to issue 7's Optimizer Memory Schedules which used training duration as an optimizer evaluation axis, this paper establishes batch size as an equally important axis and uses a noisy quadratic model to show that the bias-variance tradeoff of preconditioning can cause ranking inversions. A key reference for anyone doing optimizer comparisons or scaling law fitting.

  2. Unified hyperparameter scaling laws for MoE models treat activation ratio as a multiplicative power-law factor, enabling reliable learning rate and batch size transfer across sparsity levels.

    Editor's note

    First to include the activation rate A as an independent predictive dimension in MoE hyperparameter scaling laws, proposing the unified form h*(X,A)=k X^γ A^δ, where learning rate scales with C, batch with D, and A multiplicatively modifies both. Compared to issue 3's Let's Scale Step by Step which only fixed/limited sparsity, this paper systematically validates in the 1/64 ultra-sparse regime with 1800 runs and provides held-out extrapolation, directly addressing the reliability of hyperparameter scaling law extrapolation across MoE sparsity.

  3. Power-law learning curves emerge from cumulative weighted spectral mass near zero, not individual eigenvalues, and are jointly shaped by spectrum, target, noise, and schedule.

    Editor's note

    First to reframe the impact of joint learning rate-batch size schedules on loss power laws as a three-stage 'spectrum-memory-schedule' mechanism, providing spectral criteria for component power laws (cumulative weighted spectral quality rather than per-coordinate power laws) and boundaries where schedules preserve/alter/break clean power laws, along with memory upper bounds. Compared to issue 2's Towards Joint Scaling Laws closed-form solutions for batch schedules, this paper offers rigorous theory at the spectral mechanism level and validates B/η path equivalence and cross-schedule zero-refit predictions on 300M nanoGPT.

  4. A single shared bulk-to-spike gain in MUON matches or beats fine-grained spectral shaping across 30 settings, showing useful spectral departures are surprisingly low-dimensional.

    Editor's note

    Through spectral diagnostics, finds that 94–97% of singular modes lie below the noise edge but are collectively aligned, and proposes BulkBoost using only two bands (bulk/spike) with a single shared gain, providing an upper bound on the maximum first-order improvement rate captured by two bands. Compared to issue 4's Spectral Allocation per-rank measurements, this paper proves that coarse-grained two-band reweighting can match fine-grained power-law spectral shaping, compressing spectral shaping from per-direction fine design to a low-dimensional allocation problem.

  5. Muon optimizer splits the classical edge-of-stability into two independent signals—loss balance and temporal alignment—which respond differently to batch size and learning rate during LLM pretraining.

    Editor's note

    First to decompose Muon's edge of stability into two independent signals: a conditionally loss-neutral boundary (T1) and temporal direction alignment (T2), proving that Muon breaks the classical picture where the two are coupled in GD, and provides a coherent corrected loss boundary 2ρ_b/η under stochastic mini-batches. Compared to issue 7's Beyond Quadratic Loss stability phase diagram for Adam, this paper targets Muon's polar decomposition geometry, offering a new diagnostic framework for designing learning rate and batch schedules, validated at 22M/130M/1B scales with public code.

  6. σTransfer enables zero-shot transfer of Laplace uncertainty estimates and prior precision from small to large neural networks, yielding up to 5000x speedups with negligible performance loss.

    Editor's note

    First to extend μP's hyperparameter transfer idea to prior precision selection in Laplace approximation: by constructing a prior kernel with width-normalized coordinates, the optimal prior precision λ* stabilizes with width, enabling zero-shot transfer of precision selected on small networks to large networks, and supporting decision transfer for active learning, OOD detection, and abstention. Compared to issue 8's Fast Learning Rate Transfer which rigorously characterizes learning rate transfer, this paper extends the transfer target to uncertainty estimation, providing a complete theoretical chain from prior kernel, posterior covariance, precision selection, to decision transfer, validated at 128→4096 and 1B→7B scales.

  7. Fault-tolerant foundation models

    Fault-hardened language models become more error-resilient as they scale, unlike fault-blind models, suggesting they learn good error-correcting codes that may enable energy-efficient inference on faulty hardware.

    Editor's note

    First to introduce hardware failure rate as a new scaling law axis, adding a curvature term α2 missing from standard laws for fault-tolerant transformers, and finds that fault tolerance improves with model scale, hinting at the emergence of 'good error-correcting codes.' Validated at 7.9M–930M scales with bootstrap uncertainty reported, opening a new dimension of hardware robustness for scaling law analysis, complementing numerical stability issues in low-precision training.

  8. A hyperparameter recipe makes learning-rate and batch-size scaling predictable, yielding Chinchilla-like sqrt(C) compute-optimal scaling for jet-tagging transformers on ATLAS data.

    Editor's note

    First to jointly model hyperparameter scaling laws (multiplicative power law for η*, b*∝√T, AR*∝√N) as prerequisites for compute-optimal scaling laws on HEP jet tagging tasks, and first to predict joint optimal model size, training duration, learning rate, and batch in data-limited settings (repeated data + early stopping). Compared to issue 8's ScAn-Bench evaluation of scaling law methodologies, this paper validates the consistency of three methods (training curve envelope, IsoFLOP, full-surface fitting), and the methodology is equally applicable to LLM pretraining.