RAG-QAF: Robust Adaptation and Generalization for Quantization-Aware Fine-Tuning
Abstract
Quantization-aware fine-tuning (QAF) is crucial for adapting pretrained large language models (LLMs) to specific downstream tasks on resource-constrained devices. However, existing QAF methods tend to prioritize in-distribution (ID) performance on target adaptation tasks, while paying limited attention to out-of-distribution (OOD) generalization. Although random parameter perturbations can improve generalization in full-precision fine-tuning, directly introducing such noise into QAF may degrade both ID and OOD performance. In this work, our multi-scale analysis reveals that random noise lacks explicit optimization guidance and overlooks local differences in quantization error. We further observe that the model’s global noise sensitivity evolves throughout training, making a fixed perturbation magnitude inadequate. Based on these findings, we propose RAG-QAF, a robust QAF framework for ID adaptation and OOD generalization. Specifically, **local error-guided adversarial optimization** uses quantization errors to guide the search for adversarial perturbations through gradient ascent. **Global adaptive curriculum learning strategy** progressively increases the overall perturbation budget to accommodate evolving noise sensitivity, enabling a transition from stable early optimization to increasingly challenging robust optimization. RAG-QAF can be integrated into existing QAF pipelines as a plug-and-play module. Extensive experiments on ID and OOD benchmarks demonstrate state-of-the-art performance with low additional training overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.