SEFIR: Safety Detection in Frozen Multimodal Large Language Models via Internal Evidence Fusion
Abstract
Multimodal safety detection is difficult because harmfulness may depend on the interaction between an image and a text prompt. A common solution is to align each multimodal backbone with safety-specific data, but this is costly to repeat across model families, scales, and safety taxonomies. To address this issue, we propose SEFIR (Safety Evidence Fusion from Internal Representations), a lightweight plug-in detector for frozen MLLMs. SEFIR does not rely on the final representation alone, nor does it concatenate all hidden states. Instead, it builds a compact safety evidence space from internal representations in the visual encoder, multimodal projector, and language model. It summarizes each selected component, keeps safety-discriminative coordinates through sparse diagnostic selection, and fuses the retained evidence using held-out component utility. A small prediction head is then trained on the fused representation for global safety classification. Across multiple MLLM backbones and multimodal safety benchmarks, SEFIR improves over matched safety-oriented baselines without backbone-level realignment. Further analyses show that safety evidence is unevenly distributed across internal components, and that sparse selection with held-out utility-guided fusion helps organize this evidence into an effective lightweight detector.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.