Searching the Quantized Weight Space for Adaptation: Task-Aware Joint Quantization and Low-Rank Initialization
Abstract
Quantized low-rank fine-tuning enables memory-efficient adaptation of large language models by training low-rank adapters on a frozen low-bit backbone. Since the backbone remains fixed during fine-tuning, the quantized weights selected at initialization define the starting point from which the adapter must learn. Existing methods mainly select quantized weights to reconstruct the pretrained model’s weights or activations. However, reconstruction quality before fine-tuning does not necessarily identify the best starting point for adaptation. We therefore view initialization as a search for quantized weights that perform well after adaptation. However, this poses a major challenge: the adapted model is not yet available during the calculation of the quantized initialization. To address this challenge, we propose Task-Aware Quantized Initialization (TAQI), which uses target-task output gradients measured before fine-tuning to guide this search. TAQI weights layer-wise reconstruction errors by gradient energy and mixes the resulting task-sensitive curvature with general calibration curvature. We show that the task-weighted error bounds the energy of local first-order task-loss perturbations. Using IBM Granite-4.0-H-350M at 4 bits, TAQI improves the primary metric by 0.45–1.33 points over GPTQ-intrinsic across medical, financial, and scientific question-answering benchmarks. These results highlight the promise of adaptation-oriented quantized initialization for quantized low-rank fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.