Behavior-Aligned Visual Optimization for Jailbreaking Vision-Language Models
Abstract
Gradient-based visual jailbreak attacks on vision-language models (VLMs) often optimize image perturbations for a fixed iteration budget using token-level target objectives that only indirectly capture harmful response behavior. We introduce BAVO, a behavior-aligned visual jailbreak framework that learns a lightweight behavior predictor from internal VLM representations to estimate response-level behavior. The predictor provides a differentiable behavioral signal for image optimization and enables instance-adaptive iterate selection without response generation or evaluation at every update. Across 917 samples from multimodal jailbreak benchmarks and four target VLMs, BAVO improves ELITE-based attack success rate (E-ASR) by 16.68–33.26 percentage points and HarmBench ASR by 10.25–24.64 points over affirmative-target PGD. BAVO also reduces the mean optimization steps from 300 to 32.7–161.2, yielding – wall-clock speedups. The performance improvements persist across three defenses, while response-behavior analysis shows shifts from refusal and intermediate responses toward successful responses. These results demonstrate that aligning visual optimization with response-level behavior improves both attack effectiveness and efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.