OPQ: On-Policy Quantization for Low-Bit Large Language Models
Abstract
Low-bit quantization reduces the memory required to deploy large language models. However, learning on fixed sequences can miss the contexts reached by a quantized model's own predictions. We propose on-policy quantization (OPQ), a framework that learns directly on student-generated trajectories. On-policy quantization faces two main challenges: correcting output discrepancies introduced by low precision and maintaining consistent internal attention representations. To address the first, Dynamic Kullback–Leibler (KL) divergence adjusts output supervision using floating and quantized views of the student. To address the second, key/value (K/V) alignment supervises their attention memory. Our theoretical analysis explains the corrective gradients of the output objective and how K/V discrepancies affect attention outputs. With 2-bit and 3-bit weights, OPQ achieves the highest average task accuracy among the evaluated low-bit methods across two model families. Its gains are largest at 2-bit precision, reaching 9.49 percentage points over the strongest baseline. Ablations further show how student trajectories, Dynamic KL and K/V alignment contribute to capability preservation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.