acceptodds
Under review as a conference paper at ICLR 2027

OPQ: On-Policy Quantization for Low-Bit Large Language Models

Abstract

Low-bit quantization reduces the memory required to deploy large language models. However, learning on fixed sequences can miss the contexts reached by a quantized model's own predictions. We propose on-policy quantization (OPQ), a framework that learns directly on student-generated trajectories. On-policy quantization faces two main challenges: correcting output discrepancies introduced by low precision and maintaining consistent internal attention representations. To address the first, Dynamic Kullback–Leibler (KL) divergence adjusts output supervision using floating and quantized views of the student. To address the second, key/value (K/V) alignment supervises their attention memory. Our theoretical analysis explains the corrective gradients of the output objective and how K/V discrepancies affect attention outputs. With 2-bit and 3-bit weights, OPQ achieves the highest average task accuracy among the evaluated low-bit methods across two model families. Its gains are largest at 2-bit precision, reaching 9.49 percentage points over the strongest baseline. Ablations further show how student trajectories, Dynamic KL and K/V alignment contribute to capability preservation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.