Full text not available for this paper
Summary (Overview)
- SOLAR (State-driven Online Learning rAte scheduleR) is a reinforcement learning (RL)-based framework for online learning-rate (LR) scheduling in LLM pretraining, using PPO as the underlying RL algorithm.
- Rather than generating the full LR trajectory, SOLAR uses a base schedule (e.g., Cosine, WSD) as a reference and learns bounded, state-dependent residual corrections for individual parameter groups, re-anchoring to the base at every step.
- SOLAR improves final perplexity over tuned static schedules and automatic LR tuners across dense models from 60M to 1B parameters (with AdamW and Muon) and MoE settings up to 3B.
- A frozen residual policy trained on a 60M proxy transfers to larger dense scales without target PPO updates, remaining effective across a fourfold base-LR range.
- Matched controls show that base anchoring and action bounds improve a global PPO controller from 27.09 to 23.74 final PPL, while group-wise control reaches 22.87.
Introduction and Theoretical Foundation
Background and Motivation
Despite decades of progress in optimization, LR schedules for pretraining billion-parameter LLMs are still hand-crafted and fixed before training begins. Standard schedules like Warmup-Cosine-Decay (Cosine) (Loshchilov & Hutter, 2017) and Warmup-Stable-Decay (WSD) (Hu et al., 2024) remain the default due to their simplicity and reliability, but they cannot adapt to evolving optimization dynamics such as shifts in loss dynamics, gradient norms, and parameter magnitudes.
Adaptive optimizers (Adam, Muon) do not resolve this limitation: although they rescale updates at the individual parameter level, the global learning rate that governs overall update magnitude is still specified in advance.
Why Online LR Scheduling is Hard at LLM Scale
Online scheduling faces three key challenges in LLM pretraining:
- Noisy signals: Loss trends and gradient norms are high-variance, non-stationary, and only weakly informative at individual steps.
- Delayed feedback: A suboptimal decision may not trigger immediate failure but silently degrades progress over a long horizon.
- Asymmetric failure modes: Overly aggressive adjustments can cause abrupt instability within a few steps, making learned exploration brittle.
Theoretical Foundation
SOLAR formulates online LR scheduling as a sequential decision process:
At step , the scheduler observes state , outputs the group-wise LR action vector , and receives rewards .
The scheduler is trained with Proximal Policy Optimization (PPO) (Schulman et al., 2017) using a parameter-shared independent update across groups. The framework builds on the Learning to Optimize (L2O) paradigm, where the scheduler dynamically adapts based on the unfolding optimization state.
Methodology
4.1 Lightweight State Representation
SOLAR uses a minimal-variance principle: a compact set of low-variance, cheap-to-compute optimization signals.
Global state summarizes macro training phase and loss dynamics:
where is normalized training progress, is current loss, measures short-window loss fluctuation, and captures the discrepancy between short-term and long-term exponential moving averages of the loss.
Local state summarizes per-group context:
where is the current mini-batch gradient, is the previous action, is the depth index, is the trainable tensor, and measures recent changes in gradient magnitude.
4.2 Action Space: Stochastic Group-wise Residual Modulation
Stochastic group-wise residual actions. Rather than predicting a deterministic scalar LR, SOLAR samples from a factorized squashed Gaussian policy:
The policy uses group-specific means and a single learnable scalar standard deviation shared across all groups.
Residual LR modulation. The base schedule encodes the coarse LR profile; SOLAR controls a residual:
where controls the residual range, linearly warmed up from 0 to during early training. Because each multiplier is applied to the current base LR, the action cannot recursively redefine later LRs.
4.3 Reward Design
The reward combines immediate improvement, an EMA trend, and group-wise stability:
where:
- rewards step-wise loss reduction
- encourages improvements over a longer EMA horizon
- is a signed, group-wise stability shaping term comparing each group's post-update gradient norm with its recent trend
4.4 Circuit-Breaker Mechanism
To handle rare loss spikes, SOLAR adds an explicit instability penalty:
where is a large penalty coefficient and is a binary instability indicator:
with as a safety threshold. When triggered, the rollout is terminated, the scheduler performs a PPO update on the truncated trajectory, and the training loop restores the latest safe checkpoint.
4.5 Implementation
SOLAR uses a two-layer MLP actor–critic trained with PPO. The full algorithm is presented as Algorithm 1 in the paper.
Empirical Validation / Results
5.1 Main Pretraining Results
Dense Llama 2 Pretraining (C4, AdamW and Muon)
Table 1: Seed-52 final-checkpoint validation perplexity on C4
| Method | 60M | 130M | 350M | 1B |
|---|---|---|---|---|
| Base Optimizer: AdamW | ||||
| Cosine | 30.49 | 24.52 | 18.31 | 16.52 |
| AvgLR Replay | 30.68 | 25.61 | 21.69 | 20.06 |
| WSD | 29.80 | 23.96 | 18.75 | 16.29 |
| CLR | 31.78 | 28.14 | 21.43 | 20.22 |
| Blockwise LR | 31.06 | 24.38 | 18.56 | 16.74 |
| AutoLRS | 31.46 | 25.02 | 20.07 | 19.08 |
| MECHANIC | 32.14 | 26.95 | 20.60 | 19.25 |
| Prodigy | 45.27 | 27.51 | 22.14 | 20.75 |
| Schedule-Free | 30.28 | 24.49 | 18.14 | 16.45 |
| SOLAR (online) | 28.92 | 22.79 | 17.23 | 14.83 |
| SOLAR (frozen) | – | 22.41 | 17.37 | 14.63 |
| Base Optimizer: Muon | ||||
| Cosine | 29.20 | 22.55 | 16.87 | 14.36 |
| AvgLR Replay | 30.23 | 24.41 | 18.99 | 17.07 |
| WSD | 29.18 | 22.47 | 16.62 | 14.31 |
| CLR | 30.79 | 26.15 | 19.19 | 18.62 |
| AutoLRS | 30.07 | 23.63 | 18.95 | 16.46 |
| MECHANIC | 31.04 | 25.52 | 18.88 | 16.68 |
| SOLAR (online) | 28.63 | 21.99 | 16.35 | 13.79 |
| SOLAR (frozen) | – | 21.87 | 16.22 | 13.54 |
Key findings:
- At 1B, SOLAR-online changes AdamW+Cosine from 16.52 to 14.83 PPL (>10% relative improvement) and Muon+Cosine from 14.36 to 13.79 (~4% relative improvement).
- SOLAR-frozen further lowers five of six seed-52 online results by reusing the source-acquired controller.
- Schedule-Free AdamW is the strongest non-SOLAR alternative at 350M but remains above SOLAR.
MoE Pretraining
On Qwen2-MoE 1B pretrained on The Pile under AdamW:
- Final PPL improves from 9.61 to 9.34
- SOLAR uses a higher average LR while producing a lower, smoother gradient-norm trace than Cosine
- A 3B DeepSeek-V2-style MoE run improves final PPL from 10.73 to 10.38
Runtime overhead: On 1B AdamW, online SOLAR adds 1.23% to full-run wall-clock time and 1.27% to steady-state step time; frozen mode adds 0.76% and 0.81%, respectively. No main-table or MoE run triggers the Circuit-Breaker rollback.
5.2 Frozen Residual-Policy Transfer
A policy trained on five 60M pretraining trajectories of 11K updates, then frozen and applied to 130M, 350M, and 1B models:
- SOLAR-frozen improves over target-tuned Cosine and WSD at every larger scale
- Frozen reuse improves from 25.04 PPL at random initialization to 22.41 after five full-length 60M source runs
- The policy beats matched Cosine runs across 0.5×, 1×, and 2× target base LRs (all six settings), including completing both 2× runs where Cosine diverges
5.3 Mechanistic Evidence
Control structure (130M Llama 2, seeds 42 and 52):
| Controller | Final PPL |
|---|---|
| Recursive global PPO controller | 27.09 |
| + Re-anchoring and action bounds | 23.74 |
| + Group-wise control | 22.87 |
| Group-wise hypergradient control | 24.26 |
| Layer-wise GANNO adaptation | 24.95 |
Temporal adaptation: SOLAR maintains a higher average LR than Cosine, while AvgLR Replay ends at 25.61 PPL. A normalized static group profile reaches 24.19 PPL vs. 22.79 for SOLAR-online.
Module-wise redistribution: Vocabulary-facing components (token embeddings, LM head) receive larger multipliers, while internal attention and MLP projections are more constrained.
Module-specific feedback: Dense transformation matrices exhibit negative LR–GradNorm correlation (e.g., embed_tokens: -0.373, L2.up_proj: -0.365), while several post-attention LayerNorm parameters show positive correlation (e.g., L11.post_attn_ln: 0.261).
Theoretical and Practical Implications
Theoretical Contributions
-
Design decomposition: SOLAR divides the scheduling problem into (a) the base schedule's warmup–decay profile and (b) learned residual corrections. Re-anchoring prevents exploratory errors from compounding into a new schedule.
-
Bounded exploration: Stochastic residual actions with bounded range ( squashing) limit the damage of any single decision, making online RL feasible within a full LLM pretraining run.
-
Safety mechanism: The Circuit-Breaker provides an explicit failure signal that encourages the scheduler to internalize stability constraints, addressing the asymmetric failure modes of LR adjustment.
Practical Implications
-
Two operating modes: SOLAR-online learns the controller within the current pretraining run; SOLAR-frozen reuses an acquired policy without target-side PPO updates or retuning.
-
Transferability: A policy trained on a 60M proxy transfers to larger dense scales, remaining effective across a fourfold base-LR range—reducing the computational cost of learned LR control.
-
Optimizer-agnostic: SOLAR works with both AdamW and Muon, leaving the optimizer update unchanged and controlling only the LR across parameter groups.
-
Low overhead: The controller adds only ~1% to wall-clock time, making it practical for production-scale pretraining.
Conclusion
SOLAR establishes a practical route to adaptive LR control in LLM pretraining by combining:
- Base anchoring to preserve the coarse LR profile
- Bounded residual actions to limit individual policy decision effects
- Group-wise feedback to allow correction across parameter tensors
- Progress-aware reward and Circuit-Breaker for safety
The framework improves final PPL over tuned static schedules and automatic LR tuners across dense (60M–1B) and MoE (up to 3B) settings under both AdamW and Muon.
Primary limitation: Validation is limited to dense models up to 1B parameters and a supplementary 3B MoE setting. Behavior at larger scales (e.g., 7B+) remains untested and is left to future work.
Future directions: Extending SOLAR to larger model scales, testing across more diverse data corpora and training regimes, and further investigating the structural allocation patterns (module-wise redistribution and module-specific feedback) learned by the policy.
Related papers
- LLM sequential decision making under uncertainty in biochemical domains
LLMs overreact to new data and fail to explore in Bayesian optimization due to context-stickiness, a competence gap, not an intention gap, fixable by removing in-context history.
- Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons
Optimal learning-rate warmup duration scales with training horizon only at high peak learning rates, following a regime-dependent law predictable from three short runs.
- Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers
Newton–Schulz iterations expose spectral moments via free scalar reductions, enabling spectrum-adaptive routine selection that cuts polar error up to 90x and improves LLM pretraining loss.