# How Local Mixing Encodes Relative Position in Global NoPE Attention

> Hybrid architectures with local mixing layers and NoPE global attention implicitly learn relative position encodings via recency bias, enabling superior length extrapolation.

- **Source:** [arXiv](https://arxiv.org/abs/2609.38109)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/t0BZWq
- **Whiteboard:** https://picx.dev/p/t0BZWq/image

## Summary

## Summary (Overview)

- **Core finding**: Hybrid architectures that interleave local mixing layers (sliding window attention, gated linear attention) with global NoPE (No Position Encoding) attention layers implicitly learn relative position encodings through a recency bias mechanism.
- **Key mechanism**: SWA and gated linear layers induce a recency bias in the residual stream that propagates to global NoPE attention logits and is accentuated through network depth.
- **Theoretical contribution**: The paper mathematically demonstrates that SWA approximates a moving-average convolution in expectation, providing a formal basis for the recency bias.
- **Empirical validation**: The recency bias is present at initialization and strengthens during training, validated across hybrid models with SWA (both RoPE and NoPE) and KDA (Kimi's gated linear variant).
- **Contrast with pure NoPE**: Unlike models with only global NoPE attention (where positional info arises solely from the causal mask), hybrid models maintain recency bias across long sequences, enabling better length extrapolation.

---

## Introduction and Theoretical Foundation

### Background

Self-attention is inherently permutation-invariant, meaning it cannot distinguish token positions without explicit position encodings (PEs). Standard approaches include:

- **RoPE (Rotary Position Encoding)**: Encodes relative position via rotation matrices but incurs computational overhead and struggles with length extrapolation beyond training sequence length.
- **NoPE**: Removing explicit PE entirely has shown mixed results—poorer in-domain performance but better length extrapolation compared to RoPE.

### The Hybrid Architecture Shift

Recent large-scale models (e.g., Kimi K3, Gemma, Qwen) have shifted toward **hybrid architectures** that interleave local mixing layers with global attention layers. Key trends:

1. **Sliding Window Attention (SWA)**: Restricts attention to a local window, reducing computational cost.
2. **Gated Linear Attention (e.g., KDA)**: Uses linear recurrence with gating for efficient long-range mixing.
3. **NoPE in global layers**: The most recent models (e.g., Kimi K3) use NoPE in global attention layers.

### Research Gap

Prior work (Puvvada et al., 2025) found that hybrid models do *not* represent absolute position in NoPE layers, but left open whether they encode *relative* position. This paper addresses that gap.

### Theoretical Foundation: SWA as Moving Average

The central theoretical insight is that SWA acts as a **moving-average convolution in expectation**:

> SWA approximates a moving-average convolution in expectation, inducing a recency bias in its output that transfers into the residual stream.

This recency bias provides a mechanism for implicit relative position encoding that differs fundamentally from the causal-mask-only positional signal in pure NoPE models.

---

## Methodology

### Theoretical Analysis

The paper develops a mathematical framework showing:

1. **SWA output distribution**: For a window size $w$, the SWA layer output at position $i$ is approximately:
   $$h_i \approx \frac{1}{w} \sum_{j=i-w+1}^{i} x_j$$
   which is a moving average over the window.

2. **Recency bias propagation**: This moving-average structure creates a bias where more recent tokens have stronger influence. The bias propagates through the residual stream:
   $$r_i = x_i + h_i \approx x_i + \frac{1}{w} \sum_{j=i-w+1}^{i} x_j$$

3. **Attention logit selection**: The global NoPE attention logits select this recency bias:
   $$\text{logit}(q_i, k_j) = q_i^T k_j$$
   where the recency structure in keys $k_j$ (derived from the biased residual stream) encodes relative position.

### Experimental Setup

The authors validate their theory across multiple hybrid architectures:

- **SWA + NoPE global attention**: Sliding window layers interleaved with global NoPE attention.
- **SWA + RoPE global attention**: Baseline with explicit RoPE for comparison.
- **KDA + NoPE global attention**: Kimi's gated linear attention variant with NoPE global layers.

They measure:

- **Recency bias in residual stream**: How strongly recent tokens dominate the representation.
- **Attention logit patterns**: Whether attention weights encode relative position.
- **Depth analysis**: How the bias evolves across network layers.
- **Training dynamics**: How the bias changes from initialization through training.

---

## Empirical Validation / Results

### Key Findings

#### 1. Recency Bias in Residual Stream

The predicted recency bias is **present at initialization** and **strengthens throughout training**:

| Model | Bias at Init | Bias After Training |
|-------|-------------|-------------------|
| SWA + NoPE | Moderate | Strong |
| SWA + RoPE | Moderate | Strong |
| KDA + NoPE | Moderate | Strong |

#### 2. Attention Logit Patterns

The global NoPE attention logits exhibit clear **relative position encoding** patterns:

- Attention weights decay with relative distance (recency preference).
- The pattern is consistent with learned relative PE, not just causal masking.

#### 3. Depth Accentuation

The recency bias is **accentuated through depth**—deeper layers show stronger positional selectivity than shallow layers. This suggests the network progressively refines positional representations.

#### 4. Long-Sequence Maintenance

A critical contrast with pure NoPE models:

> In contrast to the implicit position encodings found in models with only global NoPE attention, in which positional information arises solely from the causal mask, the recency bias in hybrid models can be maintained across long sequences.

This is validated by showing the bias persists at sequence lengths far beyond the training window, whereas pure NoPE models lose positional discrimination.

#### 5. Comparison with Explicit PE

The implicit relative PE learned in hybrid models shows:

- **Comparable in-domain performance** to explicit RoPE.
- **Superior length extrapolation** beyond training sequence length.

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **Rethinking NoPE**: The paper challenges the assumption that explicit PEs are required. Hybrid architectures provide a natural mechanism for implicit relative position encoding.

2. **Unified framework**: The moving-average view of SWA provides a formal link between local mixing and positional information, potentially generalizable to other local operations.

3. **Distinction from causal-mask-only models**: The recency bias mechanism is fundamentally different from the positional signal in pure NoPE models (which relies solely on the causal mask), explaining why hybrid models maintain position across longer sequences.

### Practical Implications

1. **Architecture design**: The findings suggest that interleaving local mixing layers is not just a computational optimization but also a positional encoding strategy.

2. **Length extrapolation**: Hybrid NoPE models may offer a path to **indefinite length extrapolation** without the degradation seen in RoPE-based models.

3. **Simplified architecture**: Removing explicit PE from global layers reduces computational overhead and simplifies the model while maintaining performance.

4. **Kimi K3 validation**: The success of Kimi K3 with NoPE global attention is now theoretically grounded, providing guidance for future large-scale model design.

---

## Conclusion

### Main Takeaways

1. **Hybrid architectures implicitly encode relative position** through a recency bias mechanism induced by local mixing layers (SWA, KDA) and propagated to global NoPE attention layers.

2. **The mechanism is theoretically grounded**: SWA approximates a moving-average convolution, creating a structured bias that attention can select.

3. **The bias is learnable and robust**: Present at initialization, strengthened during training, and maintained across long sequences—unlike pure NoPE models.

4. **Practical validation**: The findings explain the empirical success of recent large-scale models (e.g., Kimi K3) and suggest design principles for future architectures.

### Future Directions

- Investigating whether other local operations (e.g., convolutions, pooling) also induce beneficial recency biases.
- Exploring optimal window sizes and layer interleaving patterns for maximizing positional encoding quality.
- Developing explicit formulations of the implicit PE learned by hybrid models, potentially enabling direct comparison with RoPE and other explicit methods.
- Studying whether the recency bias mechanism can be combined with explicit PEs for further gains.

The paper concludes that hybrid architectures offer a promising path toward **positional encoding that extrapolates to longer sequence lengths indefinitely**, a critical capability for scaling language models to increasingly long contexts.

---

_Markdown view of https://picx.dev/p/t0BZWq, served by PicX — AI-generated visual whiteboard summaries of research papers._
