QuantThink: Quantized Reasoning Models Can Learn When to Commit
Abstract
Large reasoning models (LRMs) achieve strong performance through extended chains of thought, but their long outputs impose substantial KV-cache and inference overhead. Although KV-cache quantization effectively reduces memory cost, we find that it can paradoxically prolong reasoning by inducing repeated post-solution verification and delaying answer generation even after a correct solution has already been reached. To address this issue, we introduce QuantThink, which learns when reasoning is sufficient to terminate. QuantThink combines (i) a sufficiency-aware trajectory formulation that identifies checkpoints for verification continuation or reasoning termination, and identifies redundant post-solution verification, and (ii) a parameter-efficient reinforcement learning method Q-GRPO that optimizes termination decisions under quantized inference. Across DeepSeek-R1-Distill-Qwen, QwQ, and Phi-4 models, five reasoning benchmarks, and three KV-cache quantizers—KVQuant, KIVI, and RotateKV—QuantThink improves pass@1 by up to 4.32 percentage points while reducing generated tokens by up to 63.1% compared with the KIVI K2V2 baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.