STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
Summary (Overview)
-
Problem: Recurrent states in Delta-rule linear attention models (e.g., Qwen3.8-27B, Kimi-Linear-48B) become a memory bottleneck during concurrent serving, yet direct low-bit quantization causes severe accuracy degradation due to error propagation through successive state updates.
-
Key Insight: Quantization error impact depends on two complementary dimensions: temporal (errors in long-lived memory persist across many decoding steps) and spatial (errors in different key rows affect model outputs differently, with state magnitudes varying along both rows and columns).
-
Proposed Method: STEPQuant combines Lifetime-aware Bit Allocation (temporal) with Key-Row-Aware Dual-Axis Fitting (spatial) to allocate precision based on error magnitude, memory lifetime, and key-row impact on output error.
-
Results: Under a nominal 6-bit budget, STEPQuant closely matches FP32-state accuracy on both models; even at 4 bits, it outperforms uniform INT8. Integrated into SGLang, 6-bit STEPQuant achieves 5.03× recurrent-state compression and reduces total serving memory by up to 68.7%.
-
Implementation: Code available at https://github.com/Dreamer-Toby/STEPQuant with optimized GPU kernels integrated into SGLang.
Introduction and Theoretical Foundation
Background
Unlike conventional softmax attention (which maintains a growing KV cache), linear attention summarizes past tokens into a fixed-size recurrent state matrix. Hybrid models like Qwen3.8-27B and Kimi-Linear-48B combine gated Delta-rule recurrent memory with standard attention.
Memory bottleneck: Although state size is fixed per request, each concurrent request requires a separate persistent state. In official SGLang deployment, the FP32 state pool of Qwen exceeds the memory footprint of its BF16 weights at 70 concurrent requests.
State Update Formulation
For a single head, the state after token is . The gated Delta-rule update is:
where controls memory retention, controls write strength, and is the head output (readout).
The two architectures differ in their retention gate:
- GDN (Qwen): scalar gate per head,
- KDA (Kimi): channel-wise gates,
Both can be unified as:
Quantization Error Propagation
With quantized states, the update becomes:
where quantizes and returns its dequantized approximation . Symmetric uniform quantization is defined as:
Methodology
Temporal Dimension: Lifetime-Aware Bit Allocation
Error Propagation Analysis
Proposition 1 (Conditional error propagation): For identical inputs and gates, let denote accumulated error and the quantization error added at step . Then:
If , , and , then .
Key findings:
- Previously accumulated error propagates through ; the retention gate attenuates it, and the Delta update reduces its component along while leaving orthogonal components unchanged
- Errors in directions rarely aligned with subsequent keys depend mainly on gate decay
- When retention is close to one, errors can persist for many decoding steps
- Empirical validation: Heads with longer gate half-lives exhibit larger accumulated state error (Spearman's )
Bit Allocation Objective
For each unit (entire head in Qwen, key row in KDA), estimate:
- : reconstruction distortion at bits
- : mean log retention over calibration tokens
The error retention factor after updates is approximated by:
The lifetime weight over steps:
Given average bit budget , select bit widths to minimize lifetime-weighted distortion:
Highest-risk units are retained in FP16 as sparse pivots.
Spatial Dimension: Key-Row-Aware Dual-Axis Fitting
Key-Row Impact on Readout Error
From Eq. (5), the readout error is:
Define , where weights the contribution of error in key row . The row-impact score is:
Empirical validation: Groups with larger produce greater perplexity degradation when quantized to INT4.
Two-Axis State Geometry
- Recurrent states exhibit large-magnitude outliers along both key rows and value columns
- Maximum-to-median RMS contrasts: 10.3× (key rows) and 19.4× (value columns)
- Outliers persist throughout decoding (98.6% of sampled states exceed 3× contrast)
Dual-Axis Fitting
Represent the updated state as:
where and are scale factors for key row and value column , and is the low-bit integer.
Row factors (accounting for both magnitude and impact):
where is the row magnitude.
Column factors (minimizing impact-weighted reconstruction error):
Kernel Implementation in SGLang
- Offline: Lifetime-Aware Bit Allocation and FP16 pivot selection (no per-token overhead)
- Fused kernel: tilewise state reconstruction + Delta update + readout in a single pass
- Overlapped execution: Key-Row-Aware Dual-Axis Fitting and packed writeback run on a separate CUDA stream
Empirical Validation / Results
Experimental Setup
- Models: Qwen3.8-27B (GDN), Kimi-Linear-48B-A3B-Instruct (KDA)
- Hardware: Four NVIDIA A800 GPUs
- Baselines: FP32 state, uniform INT4/6/8
- Calibration: 32 WikiText-2 segments (2048 tokens each)
- Benchmarks: 7 long-generation reasoning tasks, 6 short-generation understanding tasks
Long-Generation Results (Table 1, BF16 weights)
| State | Qwen Avg. | Kimi Avg. |
|---|---|---|
| FP32 | 80.60 | 61.52 |
| INT8 | 71.86 | 56.02 |
| INT6 | 45.04 | 45.70 |
| INT4 | 12.73 | 21.63 |
| STEPQuant@6bit | 80.59 | 61.47 |
| STEPQuant@4bit | 80.51 | 58.52 |
Short-Generation Results (Table 2, BF16 weights)
| State | Qwen Avg. | Kimi Avg. |
|---|---|---|
| FP32 | 87.78 | 68.36 |
| INT8 | 86.25 | 67.80 |
| INT6 | 82.49 | 61.35 |
| INT4 | 65.74 | 42.24 |
| STEPQuant@6bit | 87.54 | 68.82 |
| STEPQuant@4bit | 87.63 | 68.11 |
Compatibility with W4 Weights (Table 3)
With 4-bit AWQ weights, 6-bit STEPQuant achieves 79.27% (Qwen) and 58.62% (Kimi), only 0.05 and 0.33 points below FP32-state baselines.
Component Ablation (Table 4, Qwen, 4-bit budget)
| Variant | AIME | GPQA | LCB | Avg. |
|---|---|---|---|---|
| FP32 | 87.71 | 80.81 | 85.31 | 84.61 |
| INT4 | 0.00 | 4.04 | 7.87 | 3.97 |
| Q-Mamba@4bit | 0.00 | 6.06 | 16.87 | 7.64 |
| Spatial only | 75.21 | 70.71 | 75.92 | 73.95 |
| Temporal w/o pivots | 0.00 | 6.57 | 11.94 | 6.17 |
| Temporal only | 3.96 | 11.62 | 23.03 | 12.87 |
| STEPQuant@4bit | 86.25 | 81.57 | 86.35 | 84.72 |
Key ablation findings:
- Spatial fitting: Outperforms Q-Mamba's DSQ by 66.31 points at 4 bits (73.95% vs. 7.64%)
- Pivot protection: Protecting only 1.39% of Qwen heads in FP16 improves 4-bit average by 6.70 points
- Combined effect: STEPQuant outperforms both individual components, demonstrating complementarity
Serving Efficiency
At batch size 512 on Qwen with W4 weights:
- Total memory: Reduced from 419.73 to 131.18 GiB (68.7% reduction)
- Recurrent-state memory: 80.1% reduction (5.03× compression)
- State-update time: 65.6% reduction (2.91× faster)
Generation Length Analysis
Uniform quantization substantially increases output length (e.g., Kimi generates 63.40K tokens on AIME at 4-bit, approaching the 65536-token limit with near-zero accuracy). STEPQuant maintains output lengths close to FP32-state levels, consistent with preserved accuracy.
Theoretical and Practical Implications
Theoretical Contributions
-
Formal error propagation analysis: Proposition 1 provides a rigorous characterization of how quantization errors propagate through gated Delta-rule updates, showing that retention gates and Delta updates jointly shape error persistence.
-
Temporal-spatial decomposition: The paper establishes that quantization error impact depends on both when (memory lifetime) and where (key-row position) errors occur, providing a principled framework for state quantization.
-
Lifetime-weighted optimization: The bit allocation objective (Eq. 8) formalizes how to trade off quantization distortion against error persistence, generalizing prior single-axis approaches.
Practical Implications
-
Memory-efficient serving: STEPQuant enables sub-8-bit recurrent state quantization with negligible accuracy loss, directly addressing the concurrency-driven memory bottleneck in hybrid linear attention models.
-
Deployment-ready integration: The SGLang integration with optimized kernels demonstrates practical feasibility, including overlapped computation and packed-state storage.
-
Complementarity with weight quantization: STEPQuant remains effective with 4-bit AWQ weights, making it suitable for fully quantized deployment scenarios.
-
Comparison to prior work: STEPQuant achieves 6-bit quantization with negligible degradation, while concurrent work DAMP reports preserved accuracy at 9.9 bits per state value—demonstrating that lower precision is feasible with careful spatial-temporal design.
Conclusion
STEPQuant addresses the challenge of low-bit quantization for Delta-rule recurrent states by recognizing and exploiting two complementary error dimensions:
- Temporally, quantization errors persist according to memory lifetime, motivating Lifetime-aware Bit Allocation that assigns higher precision to units with larger and longer-lived errors
- Spatially, key rows differ in readout impact and state magnitudes vary along both axes, motivating Key-Row-Aware Dual-Axis Fitting with impact-weighted scales
The combined approach achieves near-FP32 accuracy at 6 bits and outperforms uniform INT8 at 4 bits on both Qwen3.8-27B and Kimi-Linear-48B across long- and short-generation benchmarks. With SGLang integration, STEPQuant delivers substantial memory savings (up to 68.7% total memory reduction) and faster state updates (2.91×), making it a practical solution for memory-efficient concurrent serving of hybrid linear attention models.
Future directions suggested by this work include extending the spatial-temporal quantization framework to other state-space architectures, exploring even lower bit budgets with adaptive pivoting strategies, and investigating the interaction between state quantization and other compression techniques (e.g., activation quantization, speculative decoding).
Related papers
- RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
RealCompanion, a benchmark from 10 real AI-companion relationships, shows memory is rarely needed (3.4% of messages) and no detector can reliably identify when it is.
- Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
Expanding alignment coverage via span supervision hurts cross-tokenizer distillation; strict 1:1 token alignment on a compact top-16 shared vocabulary subset outperforms all complex baselines.
- SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining
SOLAR uses reinforcement learning to learn bounded residual corrections to a base learning-rate schedule, improving LLM pretraining perplexity across dense and MoE models up to 3B parameters.