acceptodds
Under review as a conference paper at ICLR 2027

Targeted Transfer Attacks on MLLMs: What a Spectral Constraint Buys

Abstract

Targeted adversarial images crafted on open surrogate encoders can transfer to multimodal large language models (MLLMs) that the attacker never queries. The strongest such attacks spend the full budget, which leaves open how much perturbation transfer requires and what spending less buys. We study this question through spectral constraints on the perturbation. We insert three rules into one surrogate-only attack skeleton: hard rank truncation, uniform singular-value shrinkage and a gradient-informed threshold. We evaluate them in two regimes. In a high-transfer regime we read held-out encoder success, perceptual distortion and transfer to six generative MLLMs. In a stealth regime at a smaller step we read detectability under a residual-based detector calibrated without any constrained attack. The results separate the skeleton from the rule. Across five attack objectives, the skeleton with no rule reproduces the success gain of the complete configuration over its released baseline, so the rules we ran do not buy success. What a rule shifts is the trade-off at comparable success. In the stealth regime the soft rules lower LPIPS by about one third and reduce detector flagging from near saturation to roughly –, and the gradient-informed rule still reaches four open-weight MLLMs with every paired difference from the full-budget baseline covering . Controls show that neither scaling a perturbation down to the rule's amplitude nor lowering its rank alone reproduces this, and recalibrating the detector on reduced-amplitude copies of the attack it already knows recovers most of the detection, so the evasion depends on detector coverage. Detection and generative transfer are read on different victims. Spectral constraints therefore do not make transfer attacks more successful by themselves. They make transferable perturbations more efficient and less detectable under the studied detector coverage. [Code](https://anonymous.4open.science/r/TASS-0C8E)

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.