Revisiting Robustness Gains of Vision-Language Models under Label-Guided Attack
Abstract
It is generally expected that, even after applying a defense, models remain less accurate on adversarial inputs than on clean inputs. However, we observe that Directional Bias-guided Defense (DBD), a test-time defense for zero-shot CLIP, achieves high adversarial accuracy under ground-truth-guided (GT-guided) attacks, even surpassing clean zero-shot CLIP. We study this counterintuitive result and identify a critical evaluation caveat: when the attack objective explicitly uses the ground-truth label, the attack process leaves recoverable label information in the attacked feature that can be exploited by a subsequent recovery mechanism. Theoretically, we analyze this result from three aspects: (1) GT-guided attacks push the features away from the true class; (2) DBD can recover such attack direction from the transformed views; and (3) after reconstructing the attacked feature along this direction, the true class margin can exceed the margin on the clean feature, accounting for why DBD can surpass clean zero-shot CLIP. Empirically, our results across multiple datasets, attack settings, and CLIP backbones support this analysis. We vary the attack-label assignment through ground-truth, pseudo, loop, fixed, and controlled-corruption settings. DBD accuracy tracks the ground-truth label proportion and changes across these assignments. We further evaluate conditional subsets, perturbation budgets, other test-time defense pipelines, and complete-pipeline adaptive attacks. Finally, we introduce **D-Pred**, a simpler direction-only prediction method that serves as a diagnostic probe of attack-induced label information. It surpasses DBD in average accuracy under GT-guided PGD-100 across all three backbones, yet approaches guessing on clean inputs, exposing attack-dependent predictive information in the recovered direction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.