acceptodds
Under review as a conference paper at ICLR 2027

Concentration-Invariant Geodesic Fusion for Calibrated Test-Time Adaptation

Abstract

Backpropagation-free test-time adaptation (TTA) provides an efficient alternative for adapting pre-trained vision-language models under distribution shift; however, the underlying causes of miscalibration in these methods remain poorly understood. In this paper, we propose CRAFT, a calibration-aware adaptation framework that retains the efficiency of the backpropagation-free paradigm. We attribute the accuracy-calibration trade-off to a phenomenon we term Geometric Collapse: the unnormalized Euclidean fusion of textual anchors and visual prototypes departs from the unit hypersphere upon which the model was pre-trained, inflating the concentration unevenly across classes and resulting in overconfident predictions. To address this, CRAFT performs logit fusion along the geodesic (SLERP) on the unit hypersphere, which preserves the pre-trained concentration geometry, with a tuning-free instance-level weighting. Subsequently, a Fisher-Rao retraction on the probability simplex caps the Fisher-Rao concentration radius of the fused prediction at the VLM's inherent zero-shot level, mitigating overconfidence without requiring labeled data or tuned thresholds. Across 15 image recognition benchmarks encompassing cross-dataset benchmarks and natural distribution shifts, CRAFT consistently reduces the average Expected Calibration Error (ECE) while maintaining competitive accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.