# SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining

> SOLAR uses reinforcement learning to learn bounded residual corrections to a base learning-rate schedule, improving LLM pretraining perplexity across dense and MoE models up to 3B parameters.

- **Source:** [arXiv](https://arxiv.org/abs/2609.34681)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/ogAj3v
- **Whiteboard:** https://picx.dev/p/ogAj3v/image

## Summary

## Summary (Overview)

- **SOLAR** (State-driven Online Learning rAte scheduleR) is a reinforcement learning (RL)-based framework for online learning-rate (LR) scheduling in LLM pretraining, using PPO as the underlying RL algorithm.
- Rather than generating the full LR trajectory, SOLAR uses a **base schedule** (e.g., Cosine, WSD) as a reference and learns **bounded, state-dependent residual corrections** for individual parameter groups, re-anchoring to the base at every step.
- SOLAR improves final perplexity over tuned static schedules and automatic LR tuners across dense models from 60M to 1B parameters (with AdamW and Muon) and MoE settings up to 3B.
- A **frozen residual policy** trained on a 60M proxy transfers to larger dense scales without target PPO updates, remaining effective across a fourfold base-LR range.
- Matched controls show that base anchoring and action bounds improve a global PPO controller from 27.09 to 23.74 final PPL, while group-wise control reaches 22.87.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Despite decades of progress in optimization, LR schedules for pretraining billion-parameter LLMs are still **hand-crafted and fixed before training begins**. Standard schedules like **Warmup-Cosine-Decay (Cosine)** (Loshchilov & Hutter, 2017) and **Warmup-Stable-Decay (WSD)** (Hu et al., 2024) remain the default due to their simplicity and reliability, but they cannot adapt to evolving optimization dynamics such as shifts in loss dynamics, gradient norms, and parameter magnitudes.

Adaptive optimizers (Adam, Muon) do not resolve this limitation: although they rescale updates at the individual parameter level, the **global learning rate** that governs overall update magnitude is still specified in advance.

### Why Online LR Scheduling is Hard at LLM Scale

Online scheduling faces three key challenges in LLM pretraining:

1. **Noisy signals**: Loss trends and gradient norms are high-variance, non-stationary, and only weakly informative at individual steps.
2. **Delayed feedback**: A suboptimal decision may not trigger immediate failure but silently degrades progress over a long horizon.
3. **Asymmetric failure modes**: Overly aggressive adjustments can cause abrupt instability within a few steps, making learned exploration brittle.

### Theoretical Foundation

SOLAR formulates online LR scheduling as a **sequential decision process**:

> At step $t$, the scheduler observes state $s_t = \{s_{t,g}\}_{g=1}^G$, outputs the group-wise LR action vector $a_t = \{a_{t,g}\}_{g=1}^G$, and receives rewards $r_{t+1} = \{r_{t+1,g}\}_{g=1}^G$.

The scheduler is trained with **Proximal Policy Optimization (PPO)** (Schulman et al., 2017) using a parameter-shared independent update across groups. The framework builds on the **Learning to Optimize (L2O)** paradigm, where the scheduler dynamically adapts based on the unfolding optimization state.

---

## Methodology

### 4.1 Lightweight State Representation

SOLAR uses a minimal-variance principle: a compact set of low-variance, cheap-to-compute optimization signals.

**Global state** summarizes macro training phase and loss dynamics:

$$s^{\text{global}}_t = [\tau_t, \log L_t, \nu_t, \Delta\text{EMA}_t]$$

where $\tau_t$ is normalized training progress, $L_t$ is current loss, $\nu_t$ measures short-window loss fluctuation, and $\Delta\text{EMA}_t$ captures the discrepancy between short-term and long-term exponential moving averages of the loss.

**Local state** summarizes per-group context:

$$s^{\text{local}}_{t,g} = \left[\log \eta^{\text{base}}_{t,g}, \log \|\nabla_{t,g}\|, a_{t-1,g}, d_g, \log \|w_{t,g}\|, \Delta \log \|\nabla_{t,g}\|\right]$$

where $\nabla_{t,g}$ is the current mini-batch gradient, $a_{t-1,g}$ is the previous action, $d_g$ is the depth index, $w_{t,g}$ is the trainable tensor, and $\Delta \log \|\nabla_{t,g}\|$ measures recent changes in gradient magnitude.

### 4.2 Action Space: Stochastic Group-wise Residual Modulation

**Stochastic group-wise residual actions.** Rather than predicting a deterministic scalar LR, SOLAR samples from a factorized squashed Gaussian policy:

$$u_{t,g} \sim \mathcal{N}(\mu_{t,g}, \sigma^2), \quad \tilde{a}_{t,g} = \tanh(u_{t,g}), \quad a_{t,g} = \text{clip}(\tilde{a}_{t,g}, -1 + 10^{-4}, 1 - 10^{-4})$$

The policy uses group-specific means $\mu_{t,g}$ and a single learnable scalar standard deviation $\sigma$ shared across all groups.

**Residual LR modulation.** The base schedule encodes the coarse LR profile; SOLAR controls a residual:

$$\eta_{t,g} = \eta^{\text{base}}_{t,g} \cdot \exp(\alpha_t a_{t,g})$$

where $\alpha_t \geq 0$ controls the residual range, linearly warmed up from 0 to $\alpha$ during early training. Because each multiplier is applied to the current base LR, the action **cannot recursively redefine later LRs**.

### 4.3 Reward Design

The reward combines immediate improvement, an EMA trend, and group-wise stability:

$$r_{t+1,g} = r^{\text{perf}}_{t+1} + r^{\text{trend}}_{t+1} - p^{\text{stab}}_{t+1,g}$$

where:
- $r^{\text{perf}}_{t+1}$ rewards step-wise loss reduction
- $r^{\text{trend}}_{t+1}$ encourages improvements over a longer EMA horizon
- $p^{\text{stab}}_{t+1,g}$ is a signed, group-wise stability shaping term comparing each group's post-update gradient norm with its recent trend

### 4.4 Circuit-Breaker Mechanism

To handle rare loss spikes, SOLAR adds an explicit instability penalty:

$$r^{\text{final}}_{t+1,g} = r_{t+1,g} - \lambda_{cb} \psi_{t+1}$$

where $\lambda_{cb}$ is a large penalty coefficient and $\psi_{t+1}$ is a binary instability indicator:

$$\psi_{t+1} = \mathbb{1}\left[L_{t+1} > \kappa_L \cdot L^{\text{ema}}_{t+1}\right]$$

with $\kappa_L$ as a safety threshold. When triggered, the rollout is terminated, the scheduler performs a PPO update on the truncated trajectory, and the training loop restores the latest safe checkpoint.

### 4.5 Implementation

SOLAR uses a two-layer MLP actor–critic trained with PPO. The full algorithm is presented as Algorithm 1 in the paper.

---

## Empirical Validation / Results

### 5.1 Main Pretraining Results

#### Dense Llama 2 Pretraining (C4, AdamW and Muon)

**Table 1: Seed-52 final-checkpoint validation perplexity on C4**

| Method | 60M | 130M | 350M | 1B |
|---|---|---|---|---|
| **Base Optimizer: AdamW** | | | | |
| Cosine | 30.49 | 24.52 | 18.31 | 16.52 |
| AvgLR Replay | 30.68 | 25.61 | 21.69 | 20.06 |
| WSD | 29.80 | 23.96 | 18.75 | 16.29 |
| CLR | 31.78 | 28.14 | 21.43 | 20.22 |
| Blockwise LR | 31.06 | 24.38 | 18.56 | 16.74 |
| AutoLRS | 31.46 | 25.02 | 20.07 | 19.08 |
| MECHANIC | 32.14 | 26.95 | 20.60 | 19.25 |
| Prodigy | 45.27 | 27.51 | 22.14 | 20.75 |
| Schedule-Free | 30.28 | 24.49 | 18.14 | 16.45 |
| **SOLAR (online)** | **28.92** | **22.79** | **17.23** | **14.83** |
| **SOLAR (frozen)** | – | **22.41** | **17.37** | **14.63** |
| **Base Optimizer: Muon** | | | | |
| Cosine | 29.20 | 22.55 | 16.87 | 14.36 |
| AvgLR Replay | 30.23 | 24.41 | 18.99 | 17.07 |
| WSD | 29.18 | 22.47 | 16.62 | 14.31 |
| CLR | 30.79 | 26.15 | 19.19 | 18.62 |
| AutoLRS | 30.07 | 23.63 | 18.95 | 16.46 |
| MECHANIC | 31.04 | 25.52 | 18.88 | 16.68 |
| **SOLAR (online)** | **28.63** | **21.99** | **16.35** | **13.79** |
| **SOLAR (frozen)** | – | **21.87** | **16.22** | **13.54** |

Key findings:
- At 1B, SOLAR-online changes AdamW+Cosine from 16.52 to 14.83 PPL (**>10% relative improvement**) and Muon+Cosine from 14.36 to 13.79 (**~4% relative improvement**).
- SOLAR-frozen further lowers five of six seed-52 online results by reusing the source-acquired controller.
- Schedule-Free AdamW is the strongest non-SOLAR alternative at 350M but remains above SOLAR.

#### MoE Pretraining

On Qwen2-MoE 1B pretrained on The Pile under AdamW:
- Final PPL improves from **9.61 to 9.34**
- SOLAR uses a higher average LR while producing a lower, smoother gradient-norm trace than Cosine
- A 3B DeepSeek-V2-style MoE run improves final PPL from **10.73 to 10.38**

**Runtime overhead**: On 1B AdamW, online SOLAR adds 1.23% to full-run wall-clock time and 1.27% to steady-state step time; frozen mode adds 0.76% and 0.81%, respectively. No main-table or MoE run triggers the Circuit-Breaker rollback.

### 5.2 Frozen Residual-Policy Transfer

A policy trained on five 60M pretraining trajectories of 11K updates, then frozen and applied to 130M, 350M, and 1B models:
- SOLAR-frozen improves over target-tuned Cosine and WSD at every larger scale
- Frozen reuse improves from 25.04 PPL at random initialization to 22.41 after five full-length 60M source runs
- The policy beats matched Cosine runs across 0.5×, 1×, and 2× target base LRs (all six settings), including completing both 2× runs where Cosine diverges

### 5.3 Mechanistic Evidence

**Control structure** (130M Llama 2, seeds 42 and 52):

| Controller | Final PPL |
|---|---|
| Recursive global PPO controller | 27.09 |
| + Re-anchoring and action bounds | 23.74 |
| + Group-wise control | **22.87** |
| Group-wise hypergradient control | 24.26 |
| Layer-wise GANNO adaptation | 24.95 |

**Temporal adaptation**: SOLAR maintains a higher average LR than Cosine, while AvgLR Replay ends at 25.61 PPL. A normalized static group profile reaches 24.19 PPL vs. 22.79 for SOLAR-online.

**Module-wise redistribution**: Vocabulary-facing components (token embeddings, LM head) receive larger multipliers, while internal attention and MLP projections are more constrained.

**Module-specific feedback**: Dense transformation matrices exhibit **negative** LR–GradNorm correlation (e.g., embed_tokens: -0.373, L2.up_proj: -0.365), while several post-attention LayerNorm parameters show **positive** correlation (e.g., L11.post_attn_ln: 0.261).

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Design decomposition**: SOLAR divides the scheduling problem into (a) the base schedule's warmup–decay profile and (b) learned residual corrections. Re-anchoring prevents exploratory errors from compounding into a new schedule.

2. **Bounded exploration**: Stochastic residual actions with bounded range ($\tanh$ squashing) limit the damage of any single decision, making online RL feasible within a full LLM pretraining run.

3. **Safety mechanism**: The Circuit-Breaker provides an explicit failure signal that encourages the scheduler to internalize stability constraints, addressing the asymmetric failure modes of LR adjustment.

### Practical Implications

1. **Two operating modes**: SOLAR-online learns the controller within the current pretraining run; SOLAR-frozen reuses an acquired policy without target-side PPO updates or retuning.

2. **Transferability**: A policy trained on a 60M proxy transfers to larger dense scales, remaining effective across a fourfold base-LR range—reducing the computational cost of learned LR control.

3. **Optimizer-agnostic**: SOLAR works with both AdamW and Muon, leaving the optimizer update unchanged and controlling only the LR across parameter groups.

4. **Low overhead**: The controller adds only ~1% to wall-clock time, making it practical for production-scale pretraining.

---

## Conclusion

SOLAR establishes a practical route to adaptive LR control in LLM pretraining by combining:
- **Base anchoring** to preserve the coarse LR profile
- **Bounded residual actions** to limit individual policy decision effects
- **Group-wise feedback** to allow correction across parameter tensors
- **Progress-aware reward** and **Circuit-Breaker** for safety

The framework improves final PPL over tuned static schedules and automatic LR tuners across dense (60M–1B) and MoE (up to 3B) settings under both AdamW and Muon.

**Primary limitation**: Validation is limited to dense models up to 1B parameters and a supplementary 3B MoE setting. Behavior at larger scales (e.g., 7B+) remains untested and is left to future work.

**Future directions**: Extending SOLAR to larger model scales, testing across more diverse data corpora and training regimes, and further investigating the structural allocation patterns (module-wise redistribution and module-specific feedback) learned by the policy.

---

_Markdown view of https://picx.dev/p/ogAj3v, served by PicX — AI-generated visual whiteboard summaries of research papers._
