Full text not available for this paper

Summary

  • Core finding: Positional encoding (PE) choice systematically shapes whether language models rely on positional or semantic retrieval mechanisms for in-context retrieval.
  • Key discovery: Across 22 open-weight models spanning 8 families, standard RoPE models primarily use positional retrieval (rpos≥0.59r_{pos} \geq 0.59), while PE hybrids (e.g., SWA NoPE) shift toward semantic retrieval (rpos≤0.53r_{pos} \leq 0.53).
  • Mechanistic explanation: Confining positional encoding to local sliding-window layers degrades the model's representation of ordering IDs (positional keys), making positional retrieval less reliable and driving the shift toward semantic mechanisms.
  • Behavioral trade-off: SWA NoPE's long-context gains are not uniform—it improves on multiple-target retrieval and QA tasks but degrades when distinguishing between semantically similar competing keys.
  • Contribution: This work bridges architecture design and mechanistic interpretability, showing that retrieval strategy shifts better explain long-context performance differences than uniform improvements.

Introduction and Theoretical Foundation

Background and Motivation

Modern language models increasingly adopt hybrid architectures that vary attention span and positional encoding across layers. The most prominent example is SWA NoPE (Puvvada et al., 2025; Yang et al., 2025b), which combines:

  • Local sliding-window attention (SWA) layers with RoPE
  • Global attention layers without positional encoding (NoPE)

While prior work reported that SWA NoPE outperforms RoPE on long-context benchmarks, the underlying mechanism for this improvement remained unclear. This paper asks whether PE choice changes how models retrieve information, and whether these mechanistic differences can explain long-context retrieval performance.

Theoretical Foundation: Three Retrieval Mechanisms

Building on prior interpretability work (Gur-Arieh et al., 2026), the paper identifies three core mechanisms for in-context retrieval:

  1. Positional Mechanism (P): The model stores the position of the queried item and retrieves from that position. This relies on ordering IDs (OIs)—representations of an entity's ordinal position in a list.
  2. Lexical Mechanism (L): The model stores the queried semantic key (e.g., "hat") and retrieves the value associated with that key.
  3. Reflexive Mechanism (R): The model stores a self-referential pointer to the answer and retrieves it directly.

The paper groups lexical and reflexive as semantic mechanisms (both retrieve based on entity content) and defines the positional ratio:

rpos=PL+Rr_{pos} = \frac{P}{L + R}

Higher rposr_{pos} indicates positional dominance; lower indicates semantic dominance.

Counterfactual Patching Framework

The paper uses counterfactual activation patching to disentangle mechanisms. Given paired original and counterfactual prompts (Figure 2), patching activations at a specific layer produces different predictions depending on which mechanism dominates:

  • Positional mechanism → retrieves from the position indicated by the counterfactual
  • Lexical mechanism → retrieves using the counterfactual's semantic key
  • Reflexive mechanism → returns the self-referential pointer directly

By repeating this over many samples, the model's mechanism allocation (fraction of samples classified as P, L, R, unknown, or no effect) is estimated.

Methodology

Model Battery (Section 4.1)

The paper evaluates 22 instruction-tuned models across 8 families:

Architecture ClassModels
RoPE (13 models)Gemma-2 {2B, 9B, 27B}, Llama-3.1 {8B, 70B}, Qwen-2.5 {3B, 7B, 32B, 72B}, Qwen-3 {1.7B, 4B, 8B, 14B}
PE Hybrid (9 models)Gemma-3 {4B, 12B, 27B}, Gemma-4 {2B, 4B, 12B, 31B}, Command R7B, SmolLM3-3B

All models are evaluated on 29 entity binding tasks (from Gur-Arieh et al., 2026) with 1000 samples per task, using counterfactual patching at a model-specific retrieval layer.

Controlled Pre-training Ablation (Section 4.2)

To isolate the effect of PE choice, the paper uses 665M-parameter checkpoints from Qiao et al. (2026) with three architectures:

  • RoPE: Global attention, RoPE in all layers
  • SWA RoPE: Interleaved local/global attention (1:1 ratio, window=128), RoPE in all layers
  • SWA NoPE: Same attention pattern, but RoPE only in local layers; global layers use NoPE

Each is trained at 16K context (100B tokens) and extended to 32K (5B more tokens), totaling six checkpoints. Mechanism analysis uses 10K samples on the Boxes task at layer 15.

Linear Probing (Section 4.3)

To test whether PE choice affects ordering ID (OI) formation, the paper trains linear probes on the residual stream at layer 15:

  • Ordering ID probe: Predicts the entity's position (1–20) in the list (chance = 5%)
  • Semantic key probe (control): Predicts the object bound to each box (chance ≈ 1.2%)

Probes use L2-regularized logistic regression on 1280-dimensional activations, trained on 20,000 samples and tested on a disjoint 20,000, with 8 independent runs.

Long-Context Evaluation (Section 5)

The paper evaluates all six checkpoints on:

  • RULER (13 tasks): NIAH (single/multi-key, multi-value, multi-query), QA, variable tracking, aggregation
  • Confusable-key task (novel): Four needles with keys of four hyphenated words; varies shared words (0/4 to 3/4) to increase semantic confusability

Empirical Validation / Results

1. PE Hybrids Shift Toward Semantic Mechanisms

Figure 3 shows a complete separation between architecture classes:

MetricRoPE ModelsPE Hybrids
rposr_{pos} range≥ 0.59 (Llama-3.1 8B)≤ 0.53 (Command R7B)
Positional shareUp to 65% (Qwen2.5 72B)Down to 22% (Gemma-4 2B)
Dominant mechanismPositional for allLexical or Reflexive for all

The positional mechanism is the largest share among P, L, R for every RoPE model, while every PE hybrid has either lexical or reflexive dominant.

2. SWA Alone Does Not Cause the Shift

Table 2 shows mechanism allocation for the controlled ablation:

ArchitecturePositionalLexicalReflexiveUnknownrposr_{pos}
16K
RoPE17.656.61.624.10.30
SWA RoPE19.446.62.131.90.40
SWA NoPE10.068.61.420.00.14
32K
RoPE15.960.61.422.10.26
SWA RoPE16.257.51.824.60.27
SWA NoPE9.571.91.317.30.13

Key finding: SWA RoPE does not shift mechanisms (even increases rposr_{pos} at 16K), while SWA NoPE reduces rposr_{pos} by roughly half. The removal of RoPE from global layers is the critical factor.

3. Confining PE Degrades Ordering ID Creation

Table 3 shows linear probe accuracy at layer 15:

ContextModelOrdering IDSemantic Key
16KRoPE94.8 (0.2)99.0 (0.2)
SWA RoPE91.6 (0.1)98.8 (0.3)
SWA NoPE82.7 (0.3)99.4 (0.1)
32KRoPE94.9 (0.2)99.1 (0.2)
SWA RoPE92.7 (0.2)99.0 (0.3)
SWA NoPE83.9 (0.3)99.5 (0.1)

SWA NoPE degrades OI probe accuracy by 12.1 points (16K) and 11.0 points (32K) vs. RoPE, while SWA RoPE degrades by only 3.2 and 2.2 points. Semantic key accuracy remains nearly unchanged. The gap between semantic and positional decodability grows from ~4 to ~16 points.

4. Long-Context Retrieval Trade-off

RULER results (Table 4):

ContextModelOverall (13)NIAH mk1VT+Aggr.NIAH restQA
16KRoPE47.191.216.956.138.7
SWA NoPE58.378.017.577.243.4
32KRoPE43.392.418.150.830.7
SWA NoPE53.982.215.669.642.0

SWA NoPE improves overall RULER but fails to beat RoPE on NIAH multi-key-1, variable tracking, and aggregation tasks.

Confusable-key results (Table 5):

ContextModel0/41/42/43/4MULTIVALUEMULTIQUERY
16KRoPE97.090.884.670.835.035.0
SWA NoPE88.273.057.452.680.885.7
32KRoPE96.891.882.667.625.230.5
SWA NoPE76.067.646.838.653.865.7

Requiring discrimination among competing keys reverses the model ranking: RoPE outperforms SWA NoPE on confusable keys, while SWA NoPE dominates on multi-item retrieval. The gap widens with semantic similarity (8.8 → 27.2 points from 0/4 to 2/4 at 16K).

Theoretical and Practical Implications

Theoretical Implications

  1. Mechanism allocation is architecture-dependent: The paper demonstrates that retrieval mechanisms are not universal but are shaped by PE choice during training. Prior mechanistic analyses (Prakash et al., 2025; Gur-Arieh et al., 2026) focused only on RoPE models and missed this architectural dependence.

  2. Positional encoding's role in ordering IDs: The finding that confined PE degrades OI formation provides a mechanistic explanation for the semantic shift. Under SWA NoPE, local layers have explicit positional info but limited context, while global layers have full context but no explicit PE—making globally consistent OIs harder to learn.

  3. Long-context gains reflect strategy shifts, not uniform improvement: The paper challenges the narrative that SWA NoPE is uniformly better at long context. Instead, the mechanism shift predicts specific strengths (multi-target retrieval) and weaknesses (competing-key discrimination).

Practical Implications

  1. Intentional architecture design: The paper argues mechanistic interpretability can serve as a tool for intentional design, helping developers choose PE configurations that shape desired downstream behavior.

  2. Trade-off awareness: Model developers adopting SWA NoPE should be aware of the retrieval trade-off—improvements on QA and multi-item retrieval come at the cost of robustness to semantically similar distractors.

  3. Avoiding long-range RoPE: SWA NoPE restricts RoPE to fixed-size windows, avoiding RoPE's theoretical limitations at long context (positional/token confusion). However, this comes with the mechanistic consequences documented here.

Conclusion

Main Takeaways

  1. PE choice → mechanism allocation: Replacing uniform RoPE with SWA NoPE systematically shifts models from positional toward semantic (lexical/reflexive) retrieval mechanisms, confirmed across 22 models and controlled ablations.

  2. Mechanism shift mechanism: Confining PE to local layers degrades ordering ID representations (probe accuracy drops ~12 points), making positional retrieval less reliable and driving the semantic shift.

  3. Mechanism allocation → long-context behavior: The shift predicts distinct strengths (multi-target NIAH, QA) and failure modes (competing-key discrimination), explaining SWA NoPE's non-uniform long-context gains.

Future Directions

  • Scaling: Training larger controlled models to verify findings at scale and exhaustively ablate design choices
  • Data distribution: Studying how training data shapes mechanism allocation
  • Linear hybrids: Comprehensive analysis of linear attention hybrids (preliminary results in Appendix E show they fall between RoPE and PE hybrid ranges, with high unknown rates)
  • Window size ablations: Testing the prediction that smaller windows strengthen the semantic shift

Limitations

  • Cost of training models from scratch limits scale and exhaustive ablations
  • Mechanism allocation depends on training data distribution
  • Linear hybrids (e.g., Qwen3.5, RecurrentGemma, Granite) show high rates of unknown classifications, suggesting either mechanisms outside the current taxonomy or scale effects that require further study

Related papers