Seeing Safety Across Languages: Disentangling Multilingual Safety from Visual-Language Failures in Multimodal Large Language Models
Abstract
Multimodal safety evaluation often treats low harmful compliance or low over- refusal as direct evidence of robust safety alignment. In multilingual settings, however, such behavior may also arise from failures in visual-text recognition, language following, or task understanding. We study this problem across En- glish, Mandarin, Cantonese, and Tibetan using image-embedded instructions eval- uated on nine multimodal large language models. Our evaluation jointly measures attack success rate (ASR), over-refusal rate (ORR), target-language reply rate (TLRR), and off-topic rate (OFF), and further introduces target-language enforce- ment, cross-lingual text–image mismatch, and neutral-image controls to separate safety behavior from multilingual perception and generation failures. We find substantial language- and model-dependent variation. In particular, Cantonese and Tibetan exhibit different relationships between jailbreak susceptibility and over-refusal, while high target-language compliance does not necessarily imply successful task execution. For several models, apparently low susceptibility to Tibetan image instructions coincides with high off-topic rates or poor reproduc- tion of neutral Tibetan image text, suggesting that apparent robustness can partly reflect visual-language grounding failures rather than stronger safety alignment. These results highlight the need to evaluate multilingual multimodal safety jointly with language compliance and task understanding, rather than relying on safety rates alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.