acceptodds
Under review as a conference paper at ICLR 2027

MAAT: Massive-Activation-Aware Transferable Adversarial Attacks

Abstract

Transfer-based adversarial attacks optimise a perturbation on a surrogate model and apply it to target models without access to them. Such attacks can include objectives defined on intermediate representations, whose values depend on how representation magnitude is distributed across coordinates. We ask whether this distribution can be exploited to improve transfer. In DINOv2-L/14 and CLIP-L/14, the representation block used by the attack contains a token whose norm exceeds that of a typical token by and times, respectively, with the squared magnitude of that token concentrated in an effective dimensionality of only and channels out of . A clean-data calibration identifies the corresponding token–channel coordinates in -% of image–channel pairs. We do not assume that these coordinates are intrinsically non-transferable. Instead, MAAT uses this structure to control which channels enter the feature objective, while leaving the task loss and update rule unchanged. This makes the intervention compatible with attacks that modify the gradient, forward pass, backward pass, or input. Across two surrogates, eight targets, and four base attacks, MAAT increases mean transfer attack success rate by about four percentage points. Matched random-channel controls with six seeds support the channel selection effect for CLIP-L/14 on all seven transfer targets; the same evidence is absent for DINOv2-L/14, where an ablation identifies the source of the difference. These results show that intermediate-activation structure can be used to shape transferable adversarial objectives, with the effect depending on the surrogate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.