Fixing the Errors That Flip Answers: Output-Weighted Low-Rank Correction for Quantized Vision-Language Models
Abstract
A small low-rank correction added to a low-bit vision-language model (VLM) can repair only part of its quantization error. We ask which error directions such a limited correction budget should repair to preserve the answers of the full-precision (FP16) teacher. Starting from the observation that flips concentrate where the teacher's answer margin is small, so that small output changes decide them, we weight the low-rank reconstruction of the quantization residual by input-activation statistics and by the second moment of an output gradient that measures how strongly each output direction moves the answer; we call this output-weighted quantization error reconstruction OQER. With the same correction storage (rank 64, bits/weight) on a 3-bit LLaVA-1.5-7B, it removes (POPE) and (MME) of the teacher disagreements, against and for activation-weighted reconstruction, and at rank 32, half the extra storage, it still exceeds activation-weighted rank 64. The gain holds on top of the MBQ and QIG quantizers, on LLaVA-1.5-13B and Qwen2.5-VL-7B. In the tested settings, retaining the structure and alignment of the output weighting improves fidelity, and several output gradients are effective.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.