acceptodds
Under review as a conference paper at ICLR 2027

From Visual Harmfulness to Refusal in Low Dimensions

Abstract

Safety defenses in VLMs often reuse textual refusal representations, but stronger control can come at the cost of benign and general multimodal behavior. We show that effective refusal need be neither text-conditioned nor high-dimensional. Visual harmfulness and refusal can both be captured in low-dimensional subspaces, with refusal concentrated in markedly fewer directions. We therefore propose a visually conditioned harmfulness-to-refusal mapping that connects the two low-dimensional spaces with only 272 trainable parameters. The mapping reduces overall attack success rate by 5.0–18.4 percentage points on MM-SafetyBench and 10.7–19.2 percentage points on HoliSafe, while leaving MME performance nearly unchanged and largely preserving benign behavior on MOSSBench. Comparing the two refusal spaces reveals a trade-off in which the textual refusal space provides stronger safety, while the visual refusal space better preserves benign behavior with comparable general multimodal capability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.