acceptodds
Under review as a conference paper at ICLR 2027

OPQAD: Recovering Capability under KV-Cache Quantization via On-Policy Quantization-Aware Distillation

Abstract

QAD is a common approach for recovering capability loss in post-training quantization (PTQ). However, naive off-policy QAD recovers only 41.5% of the loss for a 1.5B model under TurboQuant K2V3. To understand this bottleneck, we analyze the issue from two perspectives: quantization-induced reasoning degradation and the intrinsic limitations of off-policy QAD. We find that quantization creates a distributional "flattening effect" and suppresses key high-confidence tokens, while off-policy QAD suffers from exposure bias and layer-wise representation drift. Building on these insights, we propose On-Policy Quantization-Aware Distillation (OPQAD), which integrates an on-policy backbone, reverse KL divergence, layer-wise hidden-state alignment, and a gated forward KL to correct suppressed tokens. Evaluated across 1.5B–8B models and long CoT tasks (up to 32K length), OPQAD improves AIME24 math accuracy by 6.0–7.3 percentage points over off-policy QAD under K2V3, raising the 1.5B model's recovery rate to 58.7%, while also yielding modest gains under OSCAR. Trained solely on math data, OPQAD generalizes to out-of-distribution code and instruction-following tasks with 9.7% training overhead. Ablations confirm that OPQAD reduces student-prefix distribution errors, hidden-layer drift, and extreme probability inversions, validating its effectiveness for low-bit KV cache recovery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.