Inverse Affine Rectification for Low-Bit Post-Training Quantized Vision Transformers
Abstract
Low-bit post-training quantization (PTQ) enables efficient deployment of vision transformers but can leave systematic distortions in the outputs of the resulting quantized model. We introduce Inverse Affine Rectification (InvAR), a simple post-hoc output rectification method for fixed quantized models that requires no rerunning of quantization optimization. InvAR uses paired logits from the full-precision (FP) and quantized (Q) models on a small set of unlabeled calibration images. The key idea is to model quantization-induced output distortion as a forward mapping from FP logits to Q logits. Rather than directly predicting FP logits from Q logits, InvAR estimates this distortion with a classwise affine model and analytically inverts it to rectify the quantized outputs. This distinction is important because, in the presence of residual variation, the least-squares Q-to-FP predictor generally differs from the inverse of the fitted FP-to-Q affine map. Using identical checkpoints and calibration images, vanilla InvAR and direct regression yield average Top-1 accuracy changes of +1.33 and percentage points, respectively, across PTQ methods, ViT architectures, and bit widths. Direct inversion, however, can amplify estimation errors when calibration data are limited, especially for affine slopes near zero. InvAR therefore uses uncertainty-guided shrinkage to reduce classwise deviations from a shared affine reference before inversion. Once paired logits are collected, fitting is non-iterative and requires operations for calibration images and classes. At inference, InvAR adds only one multiplication and one addition per prediction output. Experiments on ImageNet-1K and COCO demonstrate overall improvements across multiple PTQ methods, ViT and CNN architectures, bit widths, calibration budgets, and vision tasks. Our code will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.