acceptodds
Under review as a conference paper at ICLR 2027

FAME: FAILURE-AWARE AND MODEL-ADAPTIVE EXPLORATION FOR JAILBREAKING MLLMS

Abstract

Existing jailbreak attacks against multimodal large language models (MLLMs) make limited use of feedback from failed attempts and generally apply the same exploration strategy across target models. Consequently, they may repeatedly explore regions similar to those associated with previous failures while overlooking model-specific vulnerability patterns. To address these limitations, we propose a Failure-Aware and Model-adaptive Exploration framework (FAME), which reformulates image-side jailbreaking as closed-loop candidate selection over a pregenerated image pool. Within each attack, FAME uses previous failures to suppress candidates in low-success regions through local neighborhood avoidance and global energy attenuation. Across attacks, it trains a lightweight proxy discriminator on historical outcomes to estimate target-specific attack potential and prioritize promising candidates for each MLLM. Experiments on four benchmarks and eight open-source and commercial MLLMs show that FAME consistently improves attack success rate over state-of-the-art baselines, achieving 9.1% higher average attack success rate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.