Deeper Is Not Better for Quantized CLIP: Early Readout Recovers Accuracy While Cutting Compute
Abstract
Readout depth is a decision variable. CLIP predicts a label from the direction of its final-layer image embedding, on the premise that every transformer block moves that direction closer to the correct class. We show that this premise fails for INT8-quantized CLIP and is already unreliable in full precision. In FP32 ViT-B/32, the cosine agreement between intermediate and final representations rises with depth but regresses at three of the eight candidate layers; under INT8, the deviation introduced by quantization grows from 0.09% of the native inter-layer change at Layer 1 to 75.64% at Layer 11. We model transformer depth as a noisy readout channel in which each block adds a semantic increment and a perturbation, and show that under this model a per-sample readout gated on the top-1/top-2 margin, with depth-dependent thresholds, is expected to beat full-depth readout once perturbation growth outpaces semantic gain. LRA-EE implements this rule on a frozen INT8 backbone: a patch-token average replaces the immature shallow [CLS] token, a learned gate trained without ground-truth labels decides when to stop, and exit layers are selected by an FP32-only maturity criterion applied identically to every backbone. Across four INT8 CLIP backbones on ImageNet-1K, accuracy never falls below the full-depth baseline while FLOPs fall by 7.9–14.5%: +3.27%p on ViT-B/32, +4.27%p on ViT-B/16, +3.16%p on EVA-CLIP, and +0.66%p on ViT-L/14. In other words, the quantized model is more accurate when it computes less. A four-quadrant decomposition attributes the gain to a Rescue Effect—8.1% of samples are correct at an early exit but wrong at full depth, against 4.9% the reverse—and an exact-route control splits it into a precision-independent component (+2.42%p) and a quantization-specific increment (+0.86±0.03%p over three seeds). The evidence is empirical: the backbone is never updated and the exit heads are calibrated on 2.3% of the ImageNet training images, so the result is post-training readout calibration of a compressed model rather than zero-shot deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.