acceptodds
Under review as a conference paper at ICLR 2027

Beyond Full Imitation: Direction-Guided Quantization-Aware On-Policy Distillation

Abstract

Microscaling floating-point (MXFP) formats have emerged as a promising low-precision representation for efficient large language model (LLM) training and inference, yet quantizing both weights and activations to 4 bits (also known as W4A4 quantization) can still severely degrade performance on complex reasoning tasks. While existing post-training quantization (PTQ) and quantization-aware distillation (QAD) methods struggle to fully recover BF16 performance under aggressive quantization, on-policy distillation (OPD) is naturally suited to this setting because it supervises the quantized student on its own generated trajectories and better matches the states encountered during inference. However, we find that full imitation of the BF16 teacher is unnecessary for effective recovery. Quantization can either weaken or strengthen preferences introduced by post-training, suggesting that not every deviation from the BF16 teacher should be corrected. Based on this insight, we introduce DQ-OPD, a Direction-guided Quantization-aware On-Policy Distillation method that selectively corrects preference-weakening deviations while preserving preference-strengthening ones. Across Qwen3-8B and Qwen3-30B-A3B under MXFP W4A4 quantization, DQ-OPD consistently outperforms all PTQ and QAD baselines. On Qwen3-30B-A3B, it narrows the average gap to BF16 to 0.71 points and even surpasses the BF16 teacher on AIME24. These results demonstrate that near-BF16 performance under W4A4 quantization can be achieved without fully imitating the BF16 teacher, highlighting the importance of deciding which quantization deviations to be corrected rather than applying teacher correction uniformly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.