acceptodds
Under review as a conference paper at ICLR 2027

Linear Separability Is Not Input Controllability: Intermediate-Layer Jailbreak on VLMs

Abstract

Linear probes reveal safety-related information in Vision-Language Models (VLMs), but linear safety separability does not necessarily translate into input-space controllability. Across three VLM families, empirical layer sweeps find that attacks targeting intermediate layers are more effective than those targeting the final layer, even when late-layer safety probes remain highly accurate. Our local analysis distinguishes how much image perturbations can move representations along a safety direction from how that movement affects downstream behavior. It shows that perfect linear separability can coexist with arbitrarily weak input reachability. We introduce Intermediate-Layer Boundary Targeting (IBT), which combines intermediate-layer image optimization with text conditioning. IBT optimizes bounded image perturbations to move representations toward targets on the side classified as safe by a linear probe. On MM-SafetyBench, IBT outperforms final-layer targeting on average across the three source models in attack success rate (ASR), stealth attack success rate (SASR), and semantic response scores. In black-box transfer from an open-weight source model to three closed-source VLMs, IBT outperforms all compared baselines in mean ASR. These findings establish that linear safety separability does not guarantee input-space controllability and identify the supervision layer as a practical design variable for representation-guided image attacks, with intermediate-layer supervision yielding stronger VLM jailbreaks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.