Beyond Translation: Evaluating Language-Conditioned Safety Misalignment in LLMs Across Cantonese and Tibetan
Abstract
Multilingual safety evaluation often relies on translated benchmarks, making it unclear whether observed outcomes reflect safety alignment or language-specific behavior. We evaluate ten large language models on Cantonese and nine on Tibetan across jailbreak and over-refusal settings, comparing translated benchmarks with language-specific datasets under both natural and forced target-language response conditions. Beyond attack success rate (ASR), we measure target-language response rate (TLRR), target-language-conditioned ASR, explicit textual overrefusal, and off-topic behavior. Our results show that aggregate ASR alone can be misleading, as low attack success may arise either from stronger safety alignment or from failures to respond in the requested language. We further find that translated and language-specific datasets can yield substantially different conclusions: ASR is consistently higher on the language-specific Cantonese benchmark, whereas the difference between translated and language-specific Tibetan benchmarks is more model-dependent. In addition, the relationship between jailbreak robustness and over-refusal varies across languages and datasets rather than forming a universal trade-off. These findings suggest that multilingual safety assessment should jointly consider safety alignment and target-language adherence, and that language-specific evaluation provides complementary insights beyond translated benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.