Where Does Intervenable Refusal Live in Multimodal LM Derivatives?
Abstract
Open-weight model hubs now ship not only chat checkpoints but a long tail of derivatives—fine-tunes, safety guards, vision–language models (VLMs), audio wrappers, and generative heads—yet residual-stream safety audits still often intervene at a single site, usually the text decoder . We ask a localization question under one shared mean / rank-1 difference-in-means (DiM) extract=intervene protocol, with matched-norm random and empty-rate controls: does causal chat attack lift concentrate at , or do perceptual encoders, projectors, and fusion paths (//) supply a comparable bypass? On the mixed-input settings we stress-test—HADES adversarial images paired with harmful text prompts, and spoken AdvBench via text-to-speech (TTS) or silence—text DiM at raises attack success rate (ASR) while the same mean DiM at // does not; attacks track harmful text rather than image or audio content. Margin-rank and leverage probes on a Vision leaf show that this recipe is extremal at and near chance at /, which pressures a weak-extractor reading of those flat cells for the DiM family we study. Catalog breadth and mapper-free base→derivative transfer when residual widths match provide supporting coverage under the same gates; generative U-Nets are a separate band under generation judges. The primary claim is the vs. non-text asymmetry; an intervention-oriented Class×site ledger records where this protocol still hooks as supporting coverage, not as a universal multimodal bypass.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.