acceptodds
Under review as a conference paper at ICLR 2027

Monitor More, Supervise Less: Cross-Modal Safety Monitoring for Omni-Models

Abstract

Omni-models can receive harmful requests through different modalities and their combinations. Yet safety monitors that detect such requests from internal representations often generalize poorly beyond the modalities on which they are trained. With modalities, a request can appear in up to nonempty modality configurations, making exhaustive safety supervision and continual updating costly. This raises a central question: Can we achieve effective cross-modal safety monitoring in a data-efficient way? We study this question by formulating three supervision regimes, which compare monitors trained with safety data from one individual modality (), every individual modality (), and every nonempty modality configuration (). Instantiating with text, image, and audio modalities, we evaluate all three regimes on two datasets across five omni-models. The performance differences between successive regimes define the modality gap and the combination gap. Empirically, the modality gap is generally larger, while much of the combination gap can be recovered with a simple strategy that retains safety supervision. To reduce the modality gap, we learn a modality-shared, safety-informed subspace while keeping safety supervision. Our method improves mean AUC by and points on the two datasets, closing 57.1% and 68.4% of the modality gap, respectively. Initialized only once, the learned subspace can be reused across new safety datasets, with strong performance in cross-dataset generalization and continual monitor updating.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.