Summary (Overview)

  • Key Finding: Four seemingly minor architectural decisions—normalization strategy (QK norm, norm placement), grouped-query attention (GQA), sliding window attention, and pretraining context length—have a compoundingly negative effect on long-context extensibility, dropping downstream performance by up to 47% when combined.
  • Critical Insight: These architectural differences are not detectable from short-context loss or validation datasets, making early prediction of long-context capability extremely difficult.
  • Resource: The authors release OlmPool, a set of 26 comparable 7B models with checkpoints before and after long-context extension, trained over 170,000 GPU hours with controlled ablations holding data, tokenizer, and extension recipe fixed.
  • Surprising Result: The Llama 3 architecture is particularly well-suited for long-context extension, outperforming Qwen 3 and Olmo 3 architectures even when trained on identical data—suggesting this is an architectural, not data-driven, phenomenon.
  • Mechanistic Analysis: Attention sink behavior and attention entropy cluster by QK norm presence, with attention sinks correlating with improved long-context performance in this setting.

Introduction and Theoretical Foundation

Background and Motivation

Pretraining large language models is prohibitively expensive, yet architectural decisions must be made long before capabilities like long-context processing can be directly observed. Context length is typically extended by modifying positional embeddings and continuing pretraining at longer context lengths during a mid-training phase at the end of pretraining (Xiong et al., 2024).

The problem is compounded by the fact that most long-context extension recipes are developed on a small set of base models—primarily Llama family models (Fu et al., 2024; Gao et al., 2025; Lu et al., 2024b; Chen et al., 2024; Peng et al., 2026). This raises the question: Is the base architecture itself a decisive factor in downstream long-context performance, even when comparing only dense transformer models?

Theoretical Foundation

The authors identify four architectural decisions that vary across major model families (Llama, Qwen, Olmo) and have explicit connections to the attention mechanism or context length:

  1. Normalization strategy (QK norm variant and norm placement)
  2. Grouped-query attention (GQA) configuration
  3. Sliding window attention (SWA) usage
  4. Pretraining context length

The key mathematical formulations for QK norm are:

Layerwise QK norm (used in Olmo 2 and 3):

Q^=QRMS(Q)γQ,K^=KRMS(K)γK(1)\hat {Q} = \frac {Q}{\mathrm{RMS} (Q)} \gamma^ {Q}, \quad \hat {K} = \frac {K}{\mathrm{RMS} (K)} \gamma^ {K}\tag{1}

Headwise QK norm (used in Qwen 3, Gemma 3, Marin 32B), applied separately to each attention head:

Q^h=QhRMS(Qh)γhQ,K^h=KhRMS(Kh)γhK(2)\hat {Q} _ {h} = \frac {Q _ {h}}{\mathrm{RMS} (Q _ {h})} \gamma_ {h} ^ {Q}, \qquad \hat {K} _ {h} = \frac {K _ {h}}{\mathrm{RMS} (K _ {h})} \gamma_ {h} ^ {K}\tag{2}

Where γQ,γK\gamma ^ { Q } , \gamma ^ { K } are learned normalization parameters per layer (or per head h=1,,Hh = 1, \ldots, H for headwise variant).


Methodology

Experimental Design

The authors perform controlled pretraining experiments to construct OlmPool:

  • Model size: 26 models in the 7-8B parameter range
  • Pretraining: 140B tokens (Chinchilla-optimal amount)
  • Context extension: Adjust RoPE theta (Xiong et al., 2023), then continue pretraining for 10B tokens on 64K context data from the Longmino mix
  • Controls: Data selection/ordering, learning rate/schedule, and tokenizer held constant across all models
  • Initialization: Reused where possible; new parameters (e.g., GQA, QK norm) re-initialized

Architectural Variations Tested

FeatureVariantsModel Families Using Each
NormalizationLayerwise QK norm, headwise QK norm, no QK norm; prenorm vs. post-sublayer-normOlmo 2/3, Qwen 3, Gemma 3, Llama 2/3
GQA4, 8, 16, or 32 KV heads (32 query heads)Llama (8), Qwen (8), full attention (32)
Sliding Window3 local layers (4096 context) per 1 global layer vs. full attentionOlmo 3, Gemma 3
Pretraining Context4096 vs. 8192 tokensPrior generation vs. current generation

Note: When GQA reduced parameters, intermediate size was adjusted to maintain comparable total parameter count, which should benefit GQA models in comparisons.

Evaluation Metrics

Three complementary measures of long-context performance:

  1. RULER (Hsieh et al., 2024): Synthetic Needle-in-a-Haystack (NIAH) style tasks
  2. HELMET (Yen et al., 2025): In-context learning, reranking, question-answering
  3. LongPPL (Fang et al., 2025): Perplexity variant requiring long-range dependencies

The authors primarily report HELMET at 32K since all three measures correlate closely.


Empirical Validation / Results

Short Context Metrics Fail to Predict Long Context Performance

The models in OlmPool range from 29.9 to 56.4 on HELMET at 32K (44.7 to 67.7 on RULER at 32K)—a dramatic spread given identical data and training setup.

Key correlation findings:

PredictorR2R^2 with HELMET 32K
Training loss at end of pretraining0.29
Loss during context extension0.06
WikiText perplexity0.39
Academic text perplexity0.01–0.20
Average in-loop benchmark scores0.17
Pre-extension HELMET 8K0.32

"Standard in-loop evaluations are not sufficient to provide signal for downstream long context extensibility."

Impact of Individual Architectural Choices

Normalization is the largest individual factor:

  • Changing Olmo 3's QK norm + post-sublayer-norm to prenorm → +6 points on HELMET
  • Adding these features to Llama 3's architecture → −3.8 points

GQA degrades performance (Figure 3):

  • Increasing KV heads beyond Llama 3's 8 improves performance
  • 32 KV heads (no GQA) outperforms all GQA configurations

Sliding window and 4K pretraining context: Modest 1-2 point average degradations individually.

Compounding Effects

The most striking finding: features compound non-additively.

  • Adding sliding window to a model without GQA: −1.1 points
  • Adding sliding window to a model with GQA: −9 points on average
  • The single worst configuration combines GQA + sliding windows + headwise QK norm

Counting features is the best predictor:

  • Number of long-context-inhibiting features present: in-sample R2=0.67R^2 = 0.67, LOO R2=0.61R^2 = 0.61
  • This outperforms even a linear regression over the four individual architectural axes

Llama 3 Architecture Is Particularly Good for Long Context

The Llama 3 architecture is among the best in the design space, even when other models are pretrained on identical data. This suggests:

  • Context extension recipes developed on Llama may not transfer to other architectures
  • Olmo 3 Base being harder to context extend than Llama 3 Base is an architectural, not data, phenomenon

Robustness Checks

Token efficiency vs. capability gap (Figure 4): Even after 50B token extensions (26% of total tokens), the worst architecture does not reach Llama's performance after just 1B tokens.

Longer pretraining runs (Figure 5): Differences persist from 70B to 2T tokens; relative performance remains consistent.

Alternate extension strategy (YaRN) (Figure 6): Two-stage YaRN extension uniformly underperforms NTK one-stage, but the same relative ranking of architectures is recovered.

Attention Behavior Analysis

Attention sinks and entropy cluster by QK norm presence (Figure 7):

  • Models with QK norm have higher attention entropy (less peaky) and weaker attention sinks
  • Counterintuitively, attention sinks correlate with improved long-context performance (R2=0.38R^2 = 0.38) in OlmPool
  • The authors suggest attention sinks are the default strategy learned by QK-norm-less transformers to compensate for excess attention

Retrieval heads: Models with QK norm place less attention on needle tokens at prefill time, but during generation, little difference in retrieval head behavior was observed—possibly because these models are too weak to reliably identify retrieval heads.


Theoretical and Practical Implications

For Model Developers

  1. Architectural choices have hidden long-context costs: Each of the four features has clear benefits (training stability for normalization, inference efficiency for GQA/SWA, pretraining throughput for shorter context), but these tradeoffs must be weighed against long-context extensibility.

  2. Short-context metrics are insufficient: Developers cannot rely on standard validation to predict long-context behavior. The authors recommend performing context extension early in pretraining (as early as 70B tokens) to detect these differences.

  3. Llama-based recipes may not transfer: The field's reliance on Llama for developing context extension recipes may not generalize to other architectures.

For the Research Community

  1. OlmPool as a resource: 26 comparable models with full checkpoints (38 per model) enable future research into early pretraining phenomena, performance prediction, and mechanistic interpretability.

  2. Open questions:

    • What is the minimal size/token budget at which these effects can be reliably measured?
    • Can better proxy metrics be devised to estimate long-context performance without performing context extension?
    • What is the mechanistic basis for the differences between OlmPool models?

Conclusion

The authors demonstrate that a series of small, individually reasonable architectural perturbations—well grounded in the literature—can result in dramatically reduced long-context capabilities. The degradation is:

  • Difficult to detect in short-context metrics
  • Detectable from context extension runs very early into pretraining
  • Compound in nature: no single feature is catastrophic, but combinations of 3+ features can drop performance by up to 47%

Each feature ablated has clear benefits—stability, efficiency, throughput—yet their combination results in unacceptable long-context extensibility. By exposing the interplay between these factors in a controlled setting, the authors hope to enable model developers to make more informed architectural choices and spur research into alternatives that better navigate these tradeoffs.

Future directions:

  1. Evaluating the minimal scale at which these effects can be reliably measured
  2. Devising better proxy metrics for short-context models to estimate long-context performance
  3. Further mechanistic understanding of differences between OlmPool models
  4. Using OlmPool's parallel runs for research into other early-pretraining phenomena

Related papers