QReS: Quantizing Recurrent Language Models Is Easier Than You Think
Abstract
Linear recurrences are becoming an increasingly common building block in large language models, replacing the growing KV cache of softmax attention with a fixed-size state that enables efficient decoding. However, the quantization stack has largely been developed around softmax Transformers, whose numerical properties and inference state differ from those of recurrent models, motivating quantization schemes tailored to recurrent structure. We study how quantizing linear recurrences differs from quantizing softmax attention along two axes. First, extra-state quantization concerns weights and activations outside the recurrent state, whose distributions are shaped indirectly by the sequence mixer during training. In controlled 1.3B models, we find that decayed recurrences develop much weaker residual outlier channels than softmax attention and are therefore easier to quantize without incoherence processing. Second, intra-state quantization concerns the recurrent state itself during decode, where repeated updates and recurrent decay create a different numerical problem from KV cache quantization. We use these insights to develop QReS, a pipeline that combines value-axis scale grouping, decay-based mixed-precision allocation and stochastic rounding for recurrent state quantization. Across Qwen-3.5 9B, Nemotron-H 8B, and Kimi-Linear 48B-A3B, QReS supports robust 8-bit state quantization and lower effective precision on shorter decode horizons. We implement QReS in vLLM and demonstrate end-to-end decode speedups.
est. 46% chance this paper gets accepted at ICLR 2027.
Recent trades
| Ago | Side | Shares | Outcome | Pricea |
|---|---|---|---|---|
| 6h | buy | 234.18 | Accept | 39.4% → 46.0% |
| 1d | buy | 280.48 | Accept | 32.0% → 39.4% |
a The traded outcome’s price before and after the fill.
All positions stay anonymous.