acceptodds
Under review as a conference paper at ICLR 2027

RoQAT: A Robust Quantization-aware Training Method for LLMs Extremely Low-bit Downstream Reasoning

Abstract

Quantization-aware training (QAT) can largely recover the capabilities of low-bit large language models (LLMs). However, we discover that although achieving nearly lossless performance on general benchmarks, there still exists a tremendous gap on downstream reasoning capabilities at extremely low-bit, especially at 2-bit. We argue that this degradation is closely related to two challenges in extremely low-bit QAT: 1) directly perform low-bit training will cause severe initial degradation. Although progressive quantization can partially stabilize training, existing methods typically ignore sensitivity differences during progression; 2) the weights may converge to the sharp regions of the loss landscape, which will limit the robustness and generalization of the quantized model. Motivated by these observations, we propose RoQAT, a robust QAT framework for preserving downstream capabilities of extremely low-bit LLMs. RoQAT first introduces sensitivity-adaptive progressive quantization (SAPQ), which considers distinct quantization sensitivity across groups. Specifically, RoQAT dynamically assigns each group a different progression rate to the target bit-width according to the block quantization sensitivity along the actual progressive training direction, which allows more sensitive groups to approach the target bit-width more conservatively. Moreover, RoQAT incorporates quantization-aware sharpness minimization (QSAM) to establish flatter loss landscapes around low-precision weights, which further enhances downstream generalization. It performs sharpness-aware optimization in the weight space during block-wise training and the scale space during end-to-end refinement. Extensive experiments on Qwen3 instruction-tuned and Qwen3.5 hybrid-attention LLMs demonstrate that RoQAT consistently outperforms strong QAT baselines on downstream benchmarks while remaining stable general performance at extremely low-bit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.