acceptodds
Under review as a conference paper at ICLR 2027

IMPROVING ADVERSARIAL TRANSFERABILITY VIA CLS–PATCH RELATIONSHIP-GUIDED TOKEN GRADIENT SCALING

Abstract

The Transferability of adversarial examples pose a practical security threat to Vi￾sion Transformers (ViTs), as examples crafted on a surrogate model may also fool unseen black-box models. Existing methods improve transferability by regular￾izing token gradients or modifying attention and intermediate features. However, intervention targets are often selected using token gradients or attention information without examining how CLSpatch interactions persist or vary across layers. Since these interactions contribute to the class representation, their strength and crosslayer variation can inform a more selective intervention. We therefore propose a method that characterizes CLSpatch relationships using token representations before and after selfattention and assesses their variation across layers. Among tokens with strong relationships, we identify those exhibiting greater variation and scale their gradients at QKV and MLP outputs during backpropagation, without altering forward activations. We randomly select 1,000 images from the ImageNet validation set for evaluation. Preliminary results indicate improved black-box transferability over the MI-FGSM baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.