acceptodds
Under review as a conference paper at ICLR 2027

Rescue Vector: Bringing Quantized Reasoning Back on Track

Abstract

Low-bit quantization substantially reduces the storage and inference costs of large language models, but can cause pronounced degradation on reasoning tasks. Existing studies have mainly characterized this degradation through task-level performance or sought to mitigate it through improved quantization methods, leaving how quantization changes the reasoning process and internal representations less understood. In this work, we compare paired responses generated by the same model in its unquantized and quantized forms. We find that the two models often follow comparable reasoning initially, but the quantized model can deviate at critical intermediate steps and consequently produce incorrect answers. We further examine their hidden representations and find that the differences between the unquantized and quantized models exhibit a consistent directional pattern across failure cases. Motivated by this finding, we propose **Rescue Vector**, which aggregates these hidden-state differences to correct the intermediate representations of the quantized model. Rescue Vector can be applied directly to a quantized model, without retraining, modifying its quantized weights, or changing the original quantization procedure. Experiments across diverse language models and reasoning benchmarks show that Rescue Vector improves average accuracy under FP4 quantization from 78.76% to 81.16%, recovering 77.2% of the performance lost after quantization. On Llama-3.2-3B-Instruct, the recovery reaches 93.2%. Further analysis shows that, in rescued cases, the corrected hidden states move closer to those of the unquantized model. Together, these findings provide a new perspective on reasoning degradation under quantization from the perspective of internal representations and suggest a practical way to mitigate it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.