acceptodds
Under review as a conference paper at ICLR 2027

MAGE: Margin-Aligned Gradient Energy for Target-Supervised Robust Adaptation of Vision-Language Models

Abstract

Vision-language models are vulnerable to adversarial perturbations, and fixed confidence-based energy objectives need not guide an attacked image toward its correct class. We introduce MAGE (Margin Aligned Gradient Energy), a target-supervised, parameter-efficient robust adaptation method that learns a scalar energy head on frozen vision-language features. For each target dataset, MAGE independently trains the head using its training split, class names, and labels, shaping the negative input-energy gradient to restore the ground-truth margin while preserving clean predictions. At deployment, the adapted head performs a few bounded test-time image updates without test-label access or persistent parameter updates. Under adaptive PGD-20 that differentiates through the complete deployment pipeline, MAGE improves over fixed ET3 in all evaluated backbone-dataset combinations and achieves 37.0% average robust accuracy in the nine-dataset comparison, while also improving average clean accuracy by 7.6 points. These results show that target-supervised energy adaptation is most useful when its test-time gradient is aligned with semantic recovery rather than confidence alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.