acceptodds
Under review as a conference paper at ICLR 2027

Safety Lost in Translation: Where Harmful Intent Disappears in Vision-Language Models

Abstract

Vision-language models (VLMs) can refuse a harmful instruction in text yet comply when the same intent is delivered through pixels. Existing work establishes this safety gap across modalities, but does not isolate the computation that separates successful visual understanding from failed refusal. We introduce SIFT (Safety Information Flow Tracing), a paired analysis of the attention message sent from each modality into the shared assistant query state. To separate harmful semantics from refusal itself, refusal control is calibrated on pairs of harmful text prompts with the same underlying goal, one refused and one complied with. We localize the Safety Translation Bottleneck (STB) using a continuous refusal logit gap, with CKA, CCA, and scalar magnitude checks as robustness controls. At the STB, visual messages retain decodable harmful semantics while losing refusal control strength. Grouped edge knockout from source tokens to the query independently peaks in the same depth interval, and bidirectional message swaps on held out examples change the final response rather than only an internal score. The same mechanism motivates SIFT-BRIDGE, a rank 8 translator that repairs only the localized message channel. On a paired bank of 1,680 MM-SafetyBench intents, we obtain 0.93 AUROC for translation failure detection, compared with 0.83 for a paired multilayer probe with matched capacity. Independent replications cover LLaVA-1.5, LLaVA-1.6, Qwen2.5-VL, and InternVL3, while FigStep, VLSBench, JailBreakV-28K, and UnsafeConcepts test transfer beyond rendered text.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.