Benchmarking Test-Time Adaptation under Attack in Vision–Language Models
Abstract
Evaluating test-time adaptation under attack requires specifying the targeted predictor, the images on which errors occur, and the cost of adaptation. We compare six adaptation methods and zero-shot classification across three vision-language backbones and fifteen datasets, using common weights, starting prompts, and input images within each task. The 1,890 evaluations cover clean inputs and five cached attacks. Their per-image predictions reveal distinctions that aggregate accuracy alone misses. MAC has higher mean accuracy than R-TPT over four attacks (52.96% versus 50.86%), but a lower fraction correct under every attack (31.63% versus 37.41%); this ordering reverses within 15 of the 45 tasks. R-TPT also takes 6.33 times as long on matched attacked inputs in an 18-task local comparison. DBD leads the cached suite, yet its above-clean gains depend strongly on the attack. A separate 100-image DTD study queries all seven predictors directly. DBD's 75.56% cached accuracy falls to 34.38% under adaptive search, but a control choosing between clean and cached inputs already yields 38.94%. Thus most of that gap removes a cached-input benefit; the remaining drop under the full attack is smaller. A matched ten-image check at five times the query budget measures further sensitivity, with substantial uncertainty. These results link method selection to attack coverage, error overlap, clean-correct retention, and measured cost, while keeping shared-input performance distinct from the bounded adaptive case study.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.