OPQAD: Recovering Capability under KV-Cache Quantization via On-Policy Quantization-Aware Distillation
Abstract
QAD is a common approach for recovering capability loss in post-training quantization (PTQ). However, naive off-policy QAD recovers only 41.5% of the loss for a 1.5B model under TurboQuant K2V3. To understand this bottleneck, we analyze the issue from two perspectives: quantization-induced reasoning degradation and the intrinsic limitations of off-policy QAD. We find that quantization creates a distributional "flattening effect" and suppresses key high-confidence tokens, while off-policy QAD suffers from exposure bias and layer-wise representation drift. Building on these insights, we propose On-Policy Quantization-Aware Distillation (OPQAD), which integrates an on-policy backbone, reverse KL divergence, layer-wise hidden-state alignment, and a gated forward KL to correct suppressed tokens. Evaluated across 1.5B–8B models and long CoT tasks (up to 32K length), OPQAD improves AIME24 math accuracy by 6.0–7.3 percentage points over off-policy QAD under K2V3, raising the 1.5B model's recovery rate to 58.7%, while also yielding modest gains under OSCAR. Trained solely on math data, OPQAD generalizes to out-of-distribution code and instruction-following tasks with 9.7% training overhead. Ablations confirm that OPQAD reduces student-prefix distribution errors, hidden-layer drift, and extreme probability inversions, validating its effectiveness for low-bit KV cache recovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.