Fine-Tuning Leaves Clues: Calibration-Free Quantization via Task-Aware Basis
Abstract
Fine-tuning foundation models has become the standard way to create domain-specific expert models. While calibration-free quantization offers an efficient deployment solution without requiring downstream data, existing methods often suffer significant performance drop at low bit-widths. These methods typically neglect the task-specific information encoded in the task vector, which is defined as the parameter difference between the fine-tuned and pre-trained weights. In this work, we propose **T**ask-aware bas**I**s to **G**uid**E** calibration-f**R**ee **Q**uantization (TIGER-Q). TIGER-Q extracts a task-aware basis from the task vector to serve as a data-free proxy for the Hessian matrix. This basis establishes a task-sensitive subspace that guides the quantization process to minimize error along dimensions critical to the target task. To efficiently solve this objective, we introduce a decoupled continuous-discrete refinement pipeline. It computes closed-form analytical solutions for continuous quantization scales, while employing a local search for discrete weights, coupled with a row-wise acceptance mechanism to maintain monotonic error reduction. Extensive experiments show that TIGER-Q achieves superior performance over state-of-the-art calibration-free methods on both language and vision models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.