Linear attention swaps the ever-growing KV cache for a fixed-size recurrent state, but "fixed" only holds per request: under concurrent serving, every request keeps its own state and memory still piles up. Worse, directly quantizing those states to low precision lets errors compound through successive state updates. A Zhejiang University team tackles this engineering dead spot with STEPQuant (https://arxiv.org/abs/2609.38169).
Errors have two dimensions: how long they live, and whom they hit
The paper first answers why uniform quantization fails. Temporally, errors in long-lived memory units persist across many decoding steps. Spatially, different key rows affect outputs very differently, while state magnitudes swing along both rows and columns. Uniform bit allocation wastes precision on unimportant units while letting high-risk ones accumulate error like a snowball.
Spend bits by error, set scales by impact
STEPQuant is a post-training quantization framework: it allocates precision by error magnitude and memory lifetime, keeps a few high-risk units in FP16, and fits key-row and value-column scales separately, prioritizing rows that matter most to readout error. Calibration happens once; packed-state kernels and asynchronous writeback keep inference efficient.
The numbers: 6-bit matches FP32
Across seven long-generation tasks with BF16 weights and quantized states, Qwen3.8-27B scores 80.60% average accuracy at FP32, drops to 71.86% under uniform INT8, and collapses to 45.04% under uniform INT6; STEPQuant@6 hits 80.59% and STEPQuant@4 still reaches 80.51%. Kimi-Linear-48B-A3B-Instruct shows the same direction: FP32 61.52%, STEPQuant@6 61.47%, STEPQuant@4 58.52%, all above uniform INT8's 56.02%. On the engineering side, integrated into SGLang (tested with SGLang 0.5.12, PyTorch 2.11.0, and Triton 3.7.1; task evaluations ran on NVIDIA A800 GPUs), the 6-bit configuration compresses recurrent states by 5.03x/5.08x (Qwen/Kimi) and cuts total serving memory by up to 68.7%.
So what
Earlier this September, work on Gated DeltaNet showed full 4-bit compression of these hybrid models is feasible; STEPQuant pushes the "how to do it well" part into reproducible engineering, with code and reproduction docs open-sourced (https://github.com/Dreamer-Toby/STEPQuant). The disclosures are notably honest: the 4/6-bit budgets are nominal, since FP16 pivots and scales add storage, and the memory accounting uses W4/AWQ weights with five prefix-state slots per request. For teams optimizing hybrid-architecture inference, the spatio-temporal analysis framework matters more than the scores themselves: quantization is no longer one-size-fits-all on a tensor, but starts by asking how long an error lives and whom it affects.