Budget-Constrained Test-Time Adaptation for Vision-Language Models: A Protocol, Diagnosis, and Lightweight Fix
Abstract
Test-time adaptation (TTA) promises label-free robustness for vision-language models (VLMs) deployed in shifting environments, yet recent audits report that accuracy gains are marginal and come at the cost of trustworthiness—and none of the existing benchmarks measures the compute budget that real deployments must respect. We introduce BUDGET-TTA, an evaluation protocol with T4-GPU-hours as a first-class cost axis, covering batch=1 streaming label shift trustworthiness (ECE, OOD detection, stability), with checkpoint-resume support for low-cost cloud quotas. Using this protocol, we audit cache-based gradient-free VLM TTA and find that calibration degrades by up to 6.2 without accuracy gains—on ImageNet-V2, a simplified TDA variant matches zero-shot accuracy (52.87 vs. 52.93) while inflating ECE 6.2 (2.65 16.50); the same pattern holds on ImageNet-R. We then propose GATE-TTA, a fully backprop-free method combining confidence gating (skip when confident), smooth cache voting, sliding-window label-shift debiasing, and zero-shot anchored calibration—interpolating the adapted distribution toward the well-calibrated zero-shot one (). GATE-TTA repairs 32–66% of the ECE degradation at negligible accuracy cost, and its zero-shot-gated cache admission (7.2% OOD admission rate, 0.73 MSP-AUROC) makes it nearly immune to the cache contamination that inflates unanchored TDA's ECE by points under paired 5% OOD injection, and adds only 0.0036 T4-hours of adaptation overhead per 10k samples—orders of magnitude cheaper than gradient-based TPT ( 185 cheaper end-to-end). Our results carry a practical message: under tight budgets, skip adaptation unless the shift is severe; if you adapt, anchor to zero-shot to keep calibration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.