SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization
Abstract
Post-training quantization has emerged as an important technique for reducing the deployment cost of vision-language models (VLMs). Unlike conventional large language models, VLMs jointly process visual and textual tokens, which differ substantially in their distributions, quantities, and semantics. Existing VLM quantization methods primarily characterize such heterogeneity at either the modality or token level. However, the number, semantic roles, and visual content of tokens vary considerably across samples, making it difficult to establish stable token-level correspondences across samples. In contrast, the channel space provides a shared coordinate system across samples, making it a more natural domain in which to accumulate stable task-sensitive structures. In this work, we employ the empirical Fisher matrix as a approximation to the local curvature of the autoregressive loss with respect to intermediate activations, and perform eigendecomposition to identify a gradient-based subspace along the channel dimension. This subspace exhibits three desirable properties: cross-sample stability, clear sensitivity separation, and consistent effects along certain directions. Based on these observations, we employ a Fisher-based curvature term to suppress excessive quantization errors along sensitive directions, while using a signed first-order term to align unavoidable quantization perturbations with loss-decreasing directions. Consequently, our method goes beyond merely reducing the overall magnitude of quantization errors, further ensuring that errors within the task-sensitive subspace are both magnitude-controlled and directionally favorable. Extensive experiments demonstrate the effectiveness of our method for low-bit rotation quantization of vision-language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.