acceptodds
Under review as a conference paper at ICLR 2027

Beyond Translation: Evaluating Language-Conditioned Safety Misalignment in LLMs Across Cantonese and Tibetan

Abstract

Multilingual safety evaluation often relies on translated benchmarks, making it unclear whether observed outcomes reflect safety alignment or language-specific behavior. We evaluate ten large language models on Cantonese and nine on Tibetan across jailbreak and over-refusal settings, comparing translated benchmarks with language-specific datasets under both natural and forced target-language response conditions. Beyond attack success rate (ASR), we measure target-language response rate (TLRR), target-language-conditioned ASR, explicit textual overrefusal, and off-topic behavior. Our results show that aggregate ASR alone can be misleading, as low attack success may arise either from stronger safety alignment or from failures to respond in the requested language. We further find that translated and language-specific datasets can yield substantially different conclusions: ASR is consistently higher on the language-specific Cantonese benchmark, whereas the difference between translated and language-specific Tibetan benchmarks is more model-dependent. In addition, the relationship between jailbreak robustness and over-refusal varies across languages and datasets rather than forming a universal trade-off. These findings suggest that multilingual safety assessment should jointly consider safety alignment and target-language adherence, and that language-specific evaluation provides complementary insights beyond translated benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.