# STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

> STEPQuant achieves near-FP32 accuracy at 6-bit recurrent state quantization via lifetime-aware bit allocation and key-row-aware dual-axis fitting, cutting serving memory by 68.7%.

- **Source:** [arXiv](https://arxiv.org/abs/2609.38169)
- **Published:** 2026-10-09
- **Permalink:** https://picx.dev/p/IiRpL1
- **Whiteboard:** https://picx.dev/p/IiRpL1/image

## Summary

# STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

## Summary (Overview)

- **Problem**: Recurrent states in Delta-rule linear attention models (e.g., Qwen3.8-27B, Kimi-Linear-48B) become a memory bottleneck during concurrent serving, yet direct low-bit quantization causes severe accuracy degradation due to error propagation through successive state updates.

- **Key Insight**: Quantization error impact depends on two complementary dimensions: **temporal** (errors in long-lived memory persist across many decoding steps) and **spatial** (errors in different key rows affect model outputs differently, with state magnitudes varying along both rows and columns).

- **Proposed Method**: STEPQuant combines **Lifetime-aware Bit Allocation** (temporal) with **Key-Row-Aware Dual-Axis Fitting** (spatial) to allocate precision based on error magnitude, memory lifetime, and key-row impact on output error.

- **Results**: Under a nominal 6-bit budget, STEPQuant closely matches FP32-state accuracy on both models; even at 4 bits, it outperforms uniform INT8. Integrated into SGLang, 6-bit STEPQuant achieves **5.03× recurrent-state compression** and reduces total serving memory by up to **68.7%**.

- **Implementation**: Code available at https://github.com/Dreamer-Toby/STEPQuant with optimized GPU kernels integrated into SGLang.

---

## Introduction and Theoretical Foundation

### Background

Unlike conventional softmax attention (which maintains a growing KV cache), **linear attention** summarizes past tokens into a fixed-size recurrent state matrix. Hybrid models like Qwen3.8-27B and Kimi-Linear-48B combine gated Delta-rule recurrent memory with standard attention.

**Memory bottleneck**: Although state size is fixed per request, each concurrent request requires a separate persistent state. In official SGLang deployment, the FP32 state pool of Qwen exceeds the memory footprint of its BF16 weights at 70 concurrent requests.

### State Update Formulation

For a single head, the state after token $t$ is $S_t \in \mathbb{R}^{d_k \times d_v}$. The gated Delta-rule update is:

$$
S_t = D_t S_{t-1} + \beta_t k_t \left(v_t^\top - k_t^\top D_t S_{t-1}\right) = (I - \beta_t k_t k_t^\top) D_t S_{t-1} + \beta_t k_t v_t^\top, \quad y_t = S_t^\top q_t \tag{1}
$$

where $D_t$ controls memory retention, $\beta_t \in [0,1]$ controls write strength, and $y_t$ is the head output (readout).

The two architectures differ in their retention gate:
- **GDN** (Qwen): scalar gate per head, $D_t = \alpha_t I$
- **KDA** (Kimi): channel-wise gates, $D_t = \text{diag}(d_{t,1}, \ldots, d_{t,d_k})$

Both can be unified as:

$$
S_t = A_t S_{t-1} + B_t, \quad \text{where} \quad A_t = (I - \beta_t k_t k_t^\top) D_t, \quad B_t = \beta_t k_t v_t^\top \tag{2}
$$

### Quantization Error Propagation

With quantized states, the update becomes:

$$
X_t = A_t \hat{S}_{t-1} + B_t, \quad \hat{y}_t = X_t^\top q_t, \quad \hat{S}_t = \mathcal{Q}_t(X_t) \tag{4}
$$

where $\mathcal{Q}_t$ quantizes $X_t$ and returns its dequantized approximation $\hat{S}_t$. Symmetric uniform quantization is defined as:

$$
\mathcal{Q}_{b,s}(x) = s \cdot \text{clip}\left(\text{round}\left(\frac{x}{s}\right), -q_b, q_b\right), \qquad q_b = 2^{b-1} - 1 \tag{3}
$$

---

## Methodology

### Temporal Dimension: Lifetime-Aware Bit Allocation

#### Error Propagation Analysis

**Proposition 1 (Conditional error propagation)**: For identical inputs and gates, let $E_t = \widehat{S}_t - S_t$ denote accumulated error and $\varepsilon_t = \mathcal{Q}_t(X_t) - X_t$ the quantization error added at step $t$. Then:

$$
E_t = (I - \beta_t k_t k_t^\top) D_t E_{t-1} + \varepsilon_t = A_t E_{t-1} + \varepsilon_t, \qquad \widehat{y}_t - y_t = E_{t-1}^\top A_t^\top q_t \tag{5}
$$

If $\|k_t\|_2 \leq 1$, $0 \leq \beta_t \leq 1$, and $0 \preceq D_t \preceq I$, then $\|A_t\|_2 \leq \|D_t\|_2 \leq 1$.

**Key findings**:
- Previously accumulated error propagates through $A_t$; the retention gate $D_t$ attenuates it, and the Delta update reduces its component along $k_t$ while leaving orthogonal components unchanged
- Errors in directions rarely aligned with subsequent keys depend mainly on gate decay
- When retention is close to one, errors can persist for many decoding steps
- **Empirical validation**: Heads with longer gate half-lives exhibit larger accumulated state error (Spearman's $\rho_S \approx 0.80$)

#### Bit Allocation Objective

For each unit $u$ (entire head in Qwen, key row in KDA), estimate:
- $d_u(b)$: reconstruction distortion at $b$ bits
- $\ell_u$: mean log retention over calibration tokens

The error retention factor after $j$ updates is approximated by:

$$
\prod_{s=1}^{j} r_{t+s,u} \approx \exp(j \ell_u) \tag{6}
$$

The lifetime weight over $H$ steps:

$$
L_u = \sum_{j=0}^{H-1} \exp(2j \ell_u) \tag{7}
$$

Given average bit budget $\bar{b}$, select bit widths to minimize lifetime-weighted distortion:

$$
\min_{b_u \in \mathcal{B}_{\bar{b}}} \sum_u L_u d_u(b_u) \quad \text{s.t.} \quad \sum_u n_u b_u \leq \bar{b} \sum_u n_u \tag{8}
$$

Highest-risk units are retained in **FP16 as sparse pivots**.

### Spatial Dimension: Key-Row-Aware Dual-Axis Fitting

#### Key-Row Impact on Readout Error

From Eq. (5), the readout error is:

$$
\Delta y_t = \hat{y}_t - y_t = E_{t-1}^\top A_t^\top q_t = \sum_i (A_t^\top q_t)_i E_{t-1,i,:}^\top \tag{9}
$$

Define $g_t = A_t^\top q_t \in \mathbb{R}^{d_k}$, where $g_{t,i}$ weights the contribution of error in key row $i$. The **row-impact score** is:

$$
\omega_i = \mathbb{E}_{\text{cal}}[g_{t,i}^2] \tag{10}
$$

**Empirical validation**: Groups with larger $\omega$ produce greater perplexity degradation when quantized to INT4.

#### Two-Axis State Geometry

- Recurrent states exhibit large-magnitude outliers along **both** key rows and value columns
- Maximum-to-median RMS contrasts: **10.3×** (key rows) and **19.4×** (value columns)
- Outliers persist throughout decoding (98.6% of sampled states exceed 3× contrast)

#### Dual-Axis Fitting

Represent the updated state $X = X_t$ as:

$$
\hat{X}_{ij} = r_i c_j z_{ij} \tag{11}
$$

where $r_i > 0$ and $c_j > 0$ are scale factors for key row $i$ and value column $j$, and $z_{ij}$ is the low-bit integer.

**Row factors** (accounting for both magnitude and impact):

$$
r_i = m_i^{1/2} w_i^{-1/2} \tag{13}
$$

where $m_i = \frac{1}{d_v} \sum_j |X_{ij}|$ is the row magnitude.

**Column factors** (minimizing impact-weighted reconstruction error):

$$
\min_{\{c_j > 0\}} \sum_{i,j} w_i^2 \left(X_{ij} - r_i c_j z_{ij}\right)^2 \tag{14}
$$

### Kernel Implementation in SGLang

- Offline: Lifetime-Aware Bit Allocation and FP16 pivot selection (no per-token overhead)
- Fused kernel: tilewise state reconstruction + Delta update + readout in a single pass
- Overlapped execution: Key-Row-Aware Dual-Axis Fitting and packed writeback run on a separate CUDA stream

---

## Empirical Validation / Results

### Experimental Setup

- **Models**: Qwen3.8-27B (GDN), Kimi-Linear-48B-A3B-Instruct (KDA)
- **Hardware**: Four NVIDIA A800 GPUs
- **Baselines**: FP32 state, uniform INT4/6/8
- **Calibration**: 32 WikiText-2 segments (2048 tokens each)
- **Benchmarks**: 7 long-generation reasoning tasks, 6 short-generation understanding tasks

### Long-Generation Results (Table 1, BF16 weights)

| State | Qwen Avg. | Kimi Avg. |
|-------|-----------|-----------|
| FP32 | 80.60 | 61.52 |
| INT8 | 71.86 | 56.02 |
| INT6 | 45.04 | 45.70 |
| INT4 | 12.73 | 21.63 |
| **STEPQuant@6bit** | **80.59** | **61.47** |
| **STEPQuant@4bit** | **80.51** | **58.52** |

### Short-Generation Results (Table 2, BF16 weights)

| State | Qwen Avg. | Kimi Avg. |
|-------|-----------|-----------|
| FP32 | 87.78 | 68.36 |
| INT8 | 86.25 | 67.80 |
| INT6 | 82.49 | 61.35 |
| INT4 | 65.74 | 42.24 |
| **STEPQuant@6bit** | **87.54** | **68.82** |
| **STEPQuant@4bit** | **87.63** | **68.11** |

### Compatibility with W4 Weights (Table 3)

With 4-bit AWQ weights, 6-bit STEPQuant achieves 79.27% (Qwen) and 58.62% (Kimi), only 0.05 and 0.33 points below FP32-state baselines.

### Component Ablation (Table 4, Qwen, 4-bit budget)

| Variant | AIME | GPQA | LCB | Avg. |
|---------|------|------|-----|------|
| FP32 | 87.71 | 80.81 | 85.31 | 84.61 |
| INT4 | 0.00 | 4.04 | 7.87 | 3.97 |
| Q-Mamba@4bit | 0.00 | 6.06 | 16.87 | 7.64 |
| Spatial only | 75.21 | 70.71 | 75.92 | 73.95 |
| Temporal w/o pivots | 0.00 | 6.57 | 11.94 | 6.17 |
| Temporal only | 3.96 | 11.62 | 23.03 | 12.87 |
| **STEPQuant@4bit** | **86.25** | **81.57** | **86.35** | **84.72** |

**Key ablation findings**:
- **Spatial fitting**: Outperforms Q-Mamba's DSQ by 66.31 points at 4 bits (73.95% vs. 7.64%)
- **Pivot protection**: Protecting only 1.39% of Qwen heads in FP16 improves 4-bit average by 6.70 points
- **Combined effect**: STEPQuant outperforms both individual components, demonstrating complementarity

### Serving Efficiency

At batch size 512 on Qwen with W4 weights:
- **Total memory**: Reduced from 419.73 to 131.18 GiB (**68.7% reduction**)
- **Recurrent-state memory**: 80.1% reduction (**5.03× compression**)
- **State-update time**: 65.6% reduction (**2.91× faster**)

### Generation Length Analysis

Uniform quantization substantially increases output length (e.g., Kimi generates 63.40K tokens on AIME at 4-bit, approaching the 65536-token limit with near-zero accuracy). STEPQuant maintains output lengths close to FP32-state levels, consistent with preserved accuracy.

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Formal error propagation analysis**: Proposition 1 provides a rigorous characterization of how quantization errors propagate through gated Delta-rule updates, showing that retention gates and Delta updates jointly shape error persistence.

2. **Temporal-spatial decomposition**: The paper establishes that quantization error impact depends on both when (memory lifetime) and where (key-row position) errors occur, providing a principled framework for state quantization.

3. **Lifetime-weighted optimization**: The bit allocation objective (Eq. 8) formalizes how to trade off quantization distortion against error persistence, generalizing prior single-axis approaches.

### Practical Implications

1. **Memory-efficient serving**: STEPQuant enables sub-8-bit recurrent state quantization with negligible accuracy loss, directly addressing the concurrency-driven memory bottleneck in hybrid linear attention models.

2. **Deployment-ready integration**: The SGLang integration with optimized kernels demonstrates practical feasibility, including overlapped computation and packed-state storage.

3. **Complementarity with weight quantization**: STEPQuant remains effective with 4-bit AWQ weights, making it suitable for fully quantized deployment scenarios.

4. **Comparison to prior work**: STEPQuant achieves 6-bit quantization with negligible degradation, while concurrent work DAMP reports preserved accuracy at 9.9 bits per state value—demonstrating that lower precision is feasible with careful spatial-temporal design.

---

## Conclusion

STEPQuant addresses the challenge of low-bit quantization for Delta-rule recurrent states by recognizing and exploiting two complementary error dimensions:

- **Temporally**, quantization errors persist according to memory lifetime, motivating **Lifetime-aware Bit Allocation** that assigns higher precision to units with larger and longer-lived errors
- **Spatially**, key rows differ in readout impact and state magnitudes vary along both axes, motivating **Key-Row-Aware Dual-Axis Fitting** with impact-weighted scales

The combined approach achieves near-FP32 accuracy at 6 bits and outperforms uniform INT8 at 4 bits on both Qwen3.8-27B and Kimi-Linear-48B across long- and short-generation benchmarks. With SGLang integration, STEPQuant delivers substantial memory savings (up to 68.7% total memory reduction) and faster state updates (2.91×), making it a practical solution for memory-efficient concurrent serving of hybrid linear attention models.

**Future directions** suggested by this work include extending the spatial-temporal quantization framework to other state-space architectures, exploring even lower bit budgets with adaptive pivoting strategies, and investigating the interaction between state quantization and other compression techniques (e.g., activation quantization, speculative decoding).

---

_Markdown view of https://picx.dev/p/IiRpL1, served by PicX — AI-generated visual whiteboard summaries of research papers._
