acceptodds
Under review as a conference paper at ICLR 2027

Where and How to Steer Representations for Jailbreak Attacks on Multimodal Large Language Models

Abstract

Jailbreak attacks can induce Multimodal Large Language Models (MLLMs) to respond to harmful requests. Although numerous jailbreak attacks have been proposed, the majority of existing methods rely on output-oriented objectives to elicit affirmative responses. Such objectives are superficial and coarse-grained, which overlook the more fundamental internal representations associated with the model's refusal behavior. Meanwhile, some high-performing attacks jointly perturb visual and textual inputs, increasing the complexity of adversarial optimization. To address these limitations, we aim to identify an internal state-oriented optimization objective and develop an effective image-only attack, with a particular focus on determining where and in which direction adversarial optimization should be performed. First, to determine the attack position, we systematically analyze the representations at different positions of MLLMs. Second, to seek a suitable optimization direction, we find and validate a refusal-associated direction that can effectively influence the model's behavior. Building on these two insights, we propose a jailbreak attack that optimizes bounded perturbations to steer the representation at the target position away from the refusal-associated direction. Experiments on three MLLMs show that our method achieves attack success rates of %, %, and % on Llama-3.2-11B, Qwen2.5-VL-7B, and MiniGPT-4-13B, respectively. Despite perturbing only the input image, our method outperforms the baselines under the joint image-and-text attack settings. These results highlight the effectiveness of leveraging internal representations to guide targeted jailbreak attacks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.