QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
Abstract
Weight-only quantization is important for compressing Large Reasoning Models (LRMs). However, existing post-training quantization (PTQ) methods are suboptimal for LRMs that are typically fine-tuned from base models, because they leave their fine-tuning traces wasted during PTQ. Inspired by the spirit of classical magnitude pruning, we study whether the magnitude of weight updates during reasoning-incentivized fine-tuning can provide valuable signals for quantizing LRMs. We hypothesize that the smallest and largest weight updates are more important than those of intermediate magnitude, a phenomenon we term "protecting both ends". Upon hypothesis validation, we introduce QuantLRM, the first method for weight quantization of LRMs that explicitly leverages fine-tuning signals. We fit simple restricted quadratic functions on weight updates to protect both ends. By multiplying the average quadratic values with the count of zero weight updates of channels, we compute channel importance that is more effective than using activation or second-order information, which guides quantization by protecting important weights. To demonstrate the effectiveness and applicability of QuantLRM, we evaluate sub-4-bit quantization performance on various fine-tuning strategies (SFT, DPO, and RLVR) over different model families (Llama, Qwen, and Olmo) on a range of reasoning benchmarks (AIME-120, FOLIO, temporal sequences, and GPQA-Diamond). QuantLRM delivers consistent gains in LRMs quantization, improving average accuracy by up to 6.55% over the best baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.