acceptodds
Under review as a conference paper at ICLR 2027

QSpan: Training-Aware Low-Rank Quantization Error Compensation for Fully Quantized LoRA Fine-tuning

Abstract

Parameter-efficient fine-tuning (PEFT) reduces adaptation costs, and quantized PEFT further improves efficiency by primarily quantizing the frozen backbone weights. However, most methods retain activations, gradients, and the LoRA branch in high precision. Fully quantizing both the backbone and LoRA branch reduces these remaining costs but substantially degrades adaptation, as the frozen backbone cannot correct quantization errors and rank-constrained adapters have limited compensation capacity. Existing methods prioritize weight-quantization errors without considering their effects on the low-precision forward computation and training objective. We introduce **QSpan**, a training-aware framework that determines which quantization-error directions each adapter should span and how rank should be distributed across layers. QSpan prioritizes directions of the weight-quantization error according to their exposure through quantized activations and sensitivity to quantized gradients. It then initializes adapters along the most consequential directions through a closed-form solution and allocates a fixed total rank budget toward layers with large fractions of uncompensated training-weighted quantization error. Across language reasoning, personalized language modeling, and text-to-image personalization, QSpan consistently outperforms fully quantized LoRA and representative quantization-aware baselines under matched NVFP4 precision and adaptation budgets. These results demonstrate that aligning limited adapter capacity with training-relevant quantization errors enables effective fully quantized PEFT.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.