Pinning Visual Paths:Preserving Visual Structure for Ultra-Low-Bit Vision Encoders with Task-Informed Reconstruction
Abstract
Vision encoders increasingly rely on low-bit inference to cut computation and memory. Post-training quantization (PTQ) therefore calibrates a low-bit student to sequentially reconstruct the block-wise activations of a full-precision teacher using a subset of images. Since uniform objectives like MSE spend limited low-bit capacity on reducing all feature errors equally, regardless of their task relevance, recent methods instead optimize the loss derived from gradient-curvature-objective pipeline: they propagate quantization perturbations through a full-precision suffix (the full-precision layers following the quantized block) to obtain exact task-loss gradients, derive second-order curvature information from these gradients, and use the resulting Taylor approximation of the task loss as the optimization objective. However, these methods mainly attribute the improvement to more accurate prediction of the task loss, leaving the reliability of the estimation-and-optimization pipeline itself largely unexamined. The curvature inferred from sampled gradients may not match the Hessian required by the local Taylor expansion, and ultra-low-bit quantization produces perturbations far beyond the range where such a local approximation holds. More fundamentally, even removing this approximation and directly optimizing the exact KL divergence can overfit both the small calibration set and the full-precision suffix, while also introducing strong cross-batch variation in the optimization signal. This motivates a different objective: incorporating task information into reconstruction while keeping updates stable and reliable. We propose Pinning Visual Paths (PVP). It uses gradient magnitudes to weight block-output errors without fitting a full endpoint task-loss model, thereby avoiding unstable endpoint-KL update directions and overfitting the task-blind part of the KL update space. It adds local constraints on patch features and the next attention block's token aggregation to address the mismatch introduced by a quantized suffix. It first rotates and rescales activation coordinates to reduce outliers, then refines weight codes and activation scales to stabilize discrete optimization within a restricted update space that limits overfitting. Across seven ImageNet encoders, PVP improves mean W2A3 Top-1 accuracy by 5.31 percentage points over SOTA methods, and remains effective for self-supervised, vision-language, and downstream models. Packed kernels provide a geometric-mean 2.39× speedup and 2.99× Inmemory reduction over BF16.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.