acceptodds
Under review as a conference paper at ICLR 2027

MoMART: Mixture of Multimodal Auto-Red-Teamers

Abstract

Multimodal large language models (MLLMs) are trained for safety across the input modalities they support, yet whether these per-modality defenses remain effective under coordinated multimodal attacks remains poorly understood. We present **MoMART** (Mixture of Multimodal Auto-Red-Teamers), an adaptive compositional red-teaming framework that decomposes harmful behaviors across text, image, audio, and video using reusable modality specialists built from atomic transformations. MoMART searches over admissible modality subsets, composes the resulting artifacts into a single query, and iteratively refines attacks through a rating-guided feedback loop. Evaluating on the HarmBench benchmark against six frontier models, we find that multimodal composition consistently outperforms the strongest constituent modality under matched query budgets. On Gemini-Omni, attack success increases from for the best single modality to when all four modalities are composed, with ASR increasing monotonically across composition steps. Robustness to composition varies substantially across model families, with Claude and GPT exhibiting stronger resistance than Qwen and Llama. Beyond aggregate ASR, behavior-level matching reveals 29 and 40 joint-only successes on Gemini-Omni and Qwen-Omni, respectively, that are not achieved by any constituent modality alone. These results show that independent modality evaluation can miss vulnerabilities that emerge only through joint use, motivating adaptive, composition-aware safety assessment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.