acceptodds
Under review as a conference paper at ICLR 2027

Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation

Abstract

Recent audio-video generators increasingly support joint conditioning on text, images, audio, and video. These capabilities also enable attacks that exploit cross-modal interactions or obscure harmful intent to bypass safeguards and induce harmful audio-video outputs. However, existing generation-safety benchmarks have not kept pace with these advances, providing limited coverage of multimodal input combinations and obscured attack intents. To address these gaps, we introduce Multi2AV-Safety, the first full-coverage red-team benchmark for multimodal-to-audio-video generation, comprising 11,024 attack instances across all 11 non-singleton T/I/A/V conditioning configurations, 4 attack-intent categories, and 5 harm categories. Our evaluation of recent state-of-the-art models, including four multimodal-conditioned audio-video generators and eight safety guards, reveals substantial vulnerabilities in both generation and safeguarding, with multimodal compositional risk and obscured attack-intent risk emerging as two complementary challenges. Guided by these findings, we introduce PerceptGuard, an omni-modal guard integrating compositional-risk and attack-intent supervision through structured risk perception learning. By jointly training rationale generation and safety classification, it learns shared risk representations that enable a safety head to make efficient predictions at inference without rationale decoding, while retaining the ability to generate explanations on demand. Across 34 safety benchmarks, PerceptGuard combines SOTA multimodal safety detection with highly competitive unimodal performance, strengthening input-side safeguards against multimodal attacks on omni models. In particular, it improves safeguarding against the above risks, achieving an overall recall of 86.06% on \bench and outperforming GuardReasoner-Omni by 14.56%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.