acceptodds
Under review as a conference paper at ICLR 2027

Where Does Intervenable Refusal Live in Multimodal LM Derivatives?

Abstract

Open-weight model hubs now ship not only chat checkpoints but a long tail of derivatives—fine-tunes, safety guards, vision–language models (VLMs), audio wrappers, and generative heads—yet residual-stream safety audits still often intervene at a single site, usually the text decoder . We ask a localization question under one shared mean / rank-1 difference-in-means (DiM) extract=intervene protocol, with matched-norm random and empty-rate controls: does causal chat attack lift concentrate at , or do perceptual encoders, projectors, and fusion paths (//) supply a comparable bypass? On the mixed-input settings we stress-test—HADES adversarial images paired with harmful text prompts, and spoken AdvBench via text-to-speech (TTS) or silence—text DiM at raises attack success rate (ASR) while the same mean DiM at // does not; attacks track harmful text rather than image or audio content. Margin-rank and leverage probes on a Vision leaf show that this recipe is extremal at and near chance at /, which pressures a weak-extractor reading of those flat cells for the DiM family we study. Catalog breadth and mapper-free base→derivative transfer when residual widths match provide supporting coverage under the same gates; generative U-Nets are a separate band under generation judges. The primary claim is the vs. non-text asymmetry; an intervention-oriented Class×site ledger records where this protocol still hooks as supporting coverage, not as a universal multimodal bypass.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.