acceptodds
Under review as a conference paper at ICLR 2027

VL-ReasonHijack: A Cross-Modal Collaborative Adversarial Attack Method for Vision-Language Model Reasoning

Abstract

Vision-language models (VLMs) have developed rapidly with an emerging capability of multimodal reasoning, which aims to generate reasoning chains that connect visual evidence with commonsense knowledge, scientific concepts, or mathematical relations before producing answers. This capability creates an under-explored robustness risk: adversarial inputs can change not only the final answer, but also the reasoning chain that appears to justify it based on misleading visual evidence. Existing multimodal attack methods mainly manipulate final answers or disrupt reasoning processes, which may cause notable problems in generated texts like answer-rationale mismatch or abrupt changes in the reasoning form. Thus, we study reasoning hijacking, where an adversarial input induces a wrong answer together with a reasoning chain that supports it. To this end, we propose VL-ReasonHijack, a cross-modal collaborative adversarial attack for VLM reasoning. Given a fake reasoning chain as the hijacking objective, VL-ReasonHijack alternates between text-side question rewriting and image-side reasoning-aware optimization. Specifically, the question rewriting component emphasizes visual details used by the fake chain while preserving the original answer for a human reader. The reasoning-aware optimization component applies key reasoning atom weighting (KRAW) with a contrastive loss to reinforce reasoning-critical phrases in the fake chain. Across three mainstream VLMs and four multimodal reasoning benchmarks, VL-ReasonHijack achieves the highest answer-level attack success rates (ASR) on all the benchmark-model pairs, while producing more textually plausible reasoning outputs as reflected by lowest detection rates in most settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.