A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
Abstract
As audio-visual multi-modal large language models (MLLMs) are increasingly deployed in safety-critical applications, understanding their cross-modal vulnerabilities is essential. To this end, we introduce Multi-Modal Typography, a systematic study examining how perturbations added to speech and on-screen text impact audio-visual reasoning of closed- and open-source models. For instance, on the WorldSense benchmark, the attack success rate (ASR) against Gemini 3.7 Flash increased from 9.53% to 43.41%, highlighting its vulnerability. We analyze the cross-modal effects of unimodal and multimodal attacks and evaluate how specific properties of the attacks impact the overall multi-modal reasoning. By applying fine-tuning as a mitigation, we lower the ASR of Qwen2.5-Omni-7B on WorldSense from 60.6% to 52.0%. Our findings across multi-modal reasoning tasks, content moderation, and offline robot task selection benchmarks establish multi-modal typography as a critical and underexplored attack strategy in multi-modal reasoning. Code and data will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.