acceptodds
Under review as a conference paper at ICLR 2027

AGI-TPT: Adaptive Geometry-Informed Test-Time Prompt Tuning for Calibrated Zero-Shot Vision-Language Models

Abstract

Test-time prompt tuning (TPT) adapts a vision–language model such as CLIP to each test image by lowering the entropy of its prediction. It improves accuracy but makes the model over-confident, and when the prompt is carried along a stream of test images it collapses. We give a mechanism: entropy minimization pulls together the classes that compete for an image. We propose AGI-TPT, a regularizer whose target adapts to the number of classes N and the feature dimension D, and whose push adapts to each pair’s current state. It pushes apart only the pairs of class text features whose cosine is above the attainable level c⋆(N,D), the smallest largest cosine that N unit vectors in D dimensions can have, and it applies no force to pairs below it. The level is computed from N and D, not tuned. For free class vectors we prove that the zeros of the loss are the maximally separated configurations, that its low-rank local minima are global, and that merged classes are never local minima (N≤2D−1). We also show that at a single TPT step, spread regularizers differ mainly in strength: at equal strength AGI-TPT coincides with O-TPT’s penalty, and a random reversal of TPT’s step of the same strength comes within 0.13 ECE of it. The geometry matters when adaptation continues. There, among the prompt-tuning methods with a single template that we compare, AGI-TPT gives the best accuracy and calibration. With ViT-B/16 on 11 datasets, online, it reaches 68.8% average accuracy at ECE 1.38; at a single step it lowers the ECE of A-TPT, the best-calibrated published method, from 2.61 to 1.81 under the same protocol, and under sustained adaptation from 2.9 (A-TPT with a drift remedy) to 1.4. On four ImageNet shifts it reaches 65.3% accuracy at ECE 1.40. It keeps more of CLIP’s class-similarity structure than any baseline, and its calibration does not drift along the stream. Code will be made publicly available upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.