acceptodds
Under review as a conference paper at ICLR 2027

Safety Alignment Under Language Shift: Harmful Compliance, Over-Refusal, and Response-Language Dynamics in Low-Resource Languages

Abstract

Large language model safety alignment is usually evaluated in high-resource lan- guages, leaving low-resource settings underexplored. We evaluate proprietary and open-source models across English, Mandarin Chinese, Cantonese, and Ti- betan using benign and borderline prompts for false refusal and AdvBench for harmful compliance. We additionally measure off-topic output, repetition, and target-language response rate, and manually verify ASR-relevant labels. For ev- ery proprietary model with paired data, harmful compliance is higher in Cantonese and Tibetan than in its English–Mandarin baseline, whereas false refusal changes heterogeneously. Under unconstrained harmful prompting, Cantonese response- language consistency decreases substantially for several models, showing that aggregate safety metrics and response-language behavior should be interpreted jointly. Explicit target-language instructions largely restore language consistency but do not consistently improve safety. Among open-source models, low ASR can reflect off-topic generation or repetition, with off-topic behavior occurring across both benign and harmful datasets. Multilingual evaluation should jointly distinguish harmful compliance, false refusal, relevance, validity, and response- language behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.