Beyond ASR: Evaluating Safety Alignment of Multimodal Large Language Models Across Modalities and Low-Resource Languages
Abstract
Safety alignment in multimodal large language models (MLLMs) may vary across languages, modalities, and underlying language capabilities. We evaluate multi- ple MLLMs on harmful and benign safety-sensitive prompts in English, Man- darin, Cantonese, and Tibetan under both direct-text and image-embedded instruc- tion settings. Beyond attack success rate (ASR), we jointly analyze over-refusal rate (ORR), target-language response rate, off-topic behavior, and visual instruc- tion recovery to distinguish genuine safety behavior from failures in low-resource language understanding or multimodal processing. Our results reveal substantial safety disparities in Cantonese and Tibetan, while showing that image embedding does not uniformly increase jailbreak risk. In particular, low ASR for Tibetan image inputs is often accompanied by high off-topic rates and poor recovery of the embedded instruction, indicating that capability failure can be mistaken for stronger safety. Cross-language image–text ablations further confirm that the lan- guage carried by the visual modality substantially affects model behavior. These findings highlight the importance of capability-aware evaluation for multilingual multimodal safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.