RTAQ: Reasoning Termination-Aware Quantization for Mixture-of-Experts Architecture
Abstract
Extreme low-bit post-training quantization is key to serving large-scale mixture-of-experts (MoE) language models, yet reasoning models quantized to around two bits often collapse on long-horizon reasoning tasks. We show that this collapse stems not from a loss of knowledge but from a failure to stop reasoning properly. When a 2-bit or near-1-bit reasoning model emits the end-of-thinking token (</think>), it terminates in every case we observe and answers nearly as accurately as its full-precision counterpart; most errors arise from reasoning traces that never close. We trace this termination decision to a small subset of routed experts (approximately 2% of all routed experts) that are consistently selected at the end-of-thinking position but are largely overlooked by standard sensitivity metrics. Building on this finding, we propose RTAQ (Reasoning-Termination-Aware Quantization), a weight-only quantization pipeline with three stages: (1) constructing a bank of activation-weighted, DP-optimal scalar codebooks for every expert and bit-width; (2) performing termination-aware bit allocation, which fixes termination experts at higher precision within the bit budget and assigns the remaining bits through knapsack optimization; and (3) applying grid-preserving recalibration, which reassigns codes using GPTQ-style error compensation and rescales expert outputs in under one second per expert. At an average bit-width of 2.0 for routed experts, RTAQ preserves over 95% of the full-precision performance on long-reasoning benchmarks such as AIME.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.