acceptodds
Under review as a conference paper at ICLR 2027

Beyond First-Order Feature Alignment: Second-Order Subspace Alignment for Adversarial Attacks on Closed-Source MLLMs

Abstract

Multimodal large language models (MLLMs) are vulnerable to targeted adversarial attacks transferred from surrogate encoders to closed-source models. Recent feature-alignment based attacks exploit transformed views and fine-grained local representations, yet remain centered on first-order feature correspondence. We observe that subspace modeling offers a natural structural perspective for capturing second-order moment structure across related representations. Accordingly, we propose the Second-Order Subspace Alignment Attack (SOSA-Attack), which aligns augmentation-induced subspace descriptors derived from second-order moments at both global and patch levels and progressively fuses the resulting subspace alignment gradients with first-order feature alignment updates. Theoretically, we characterize SOSA as second-order distribution matching, proving an exact equivalence between second-order moment discrepancy and squared polynomial-kernel MMD, with corresponding bounds for the MaxExp subspace discrepancy. Experiments across diverse MLLMs and downstream tasks show that SOSA-Attack consistently outperforms prior targeted transfer attacks and achieves state-of-the-art performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.