acceptodds
Under review as a conference paper at ICLR 2027

When Quantization Helps: Corrective Computation in Low-Bit Language Models

Abstract

Low-bit quantization reduces deployment costs by approximating high-precision weights, introducing changes that can improve some predictions. We investigate which activation changes support lower target-task loss than the high-precision reference and whether their useful structure can be shared across inputs. Experiments on Llama (GPTQ, AWQ, NF4) and Qwen (GPTQ) reveal systematic input-level improvements despite lower average performance over the full evaluation set. Improved inputs exhibit stronger loss-reducing contributions per unit activation displacement. Stronger first-order loss reduction can outweigh comparable or larger loss increases beyond the linear approximation. Targeted interventions link these changes to predictions: replacing selected quantized activations with their high-precision values can reverse naturally corrected answers; strengthening favorable changes or suppressing harmful ones repairs errors. We construct useful local adjustments from quantization-induced changes and find contrasting shared structure: a rank-four subspace retains 95.3% of their cross-entropy improvement on HANS, whereas MMLU retains 47.8% even at rank 128. Mechanism-guided activation adaptation raises accuracy above the matched high-precision model by 7.41 percentage points on HANS test data and 2.42 on MMLU development data. Its benefit over ordinary supervised adaptation is larger on HANS, consistent with its stronger shared structure. Together, these results connect the loss-reducing contributions of quantization-induced changes to their effects on predictions and relate their shared structure to the benefits of mechanism-guided adaptation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.