acceptodds
Under review as a conference paper at ICLR 2027

Jailbreaking Multimodal Large Language Models through Generative Visual Semantic Ciphers

Abstract

Multimodal large language models (MLLMs) can be induced to reconstruct harm- ful instructions from distributed visual and textual cues while failing to reliably apply safety constraints to the recovered intent. We investigate this failure mode through the Generative Visual Semantic Cipher Attack (GVSCA), a black-box framework that encodes instruction components as recognizable surrogate con- cepts within a single generated natural scene. Combined with attacker-provided semantic mappings and decoding guidance, the image induces the target model to reconstruct and respond to a harmful instruction that is not presented as a con- tiguous plaintext command. Across 11 MLLMs and 65 selected harmful instruc- tions spanning eight categories, GVSCA achieves average attack success rates of 69.76% with one query and 93.01% within ten queries, compared with 14.90% and 29.79%, respectively, for the strongest evaluated baseline. Matched-carrier experiments on three models show that replacing the generated images with tex- tual carriers reduces single-query success from 64.23% to at most 5.13%, with semantic mappings and reconstruction procedures held fixed. Joint image–text filtering with Llama Guard 4 still leaves a 55.69% single-query success rate. In a separate three-model evaluation, a system instruction requiring safety assessment of the reconstructed intent reduces single-query success from 70.39% to 20.00%, but incurs a 15.00% benign refusal rate. These findings highlight a safety gap between cross-modal semantic reconstruction and safety enforcement, motivating defenses that assess reconstructed instructions while preserving benign reasoning capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.