acceptodds
Under review as a conference paper at ICLR 2027

One Teacher, Many Tongues: Deficit-Aware Safety Distillation for Multilingual Alignment

Abstract

Multilingual safety alignment must reduce harmful responses across languages while preserving reasoning ability. Achieving this goal requires allocating on-policy safety-distillation supervision across both response positions and semantically equivalent language instances. We introduce DASD (Deficit-Aware Safety Distillation), a normalized token-level KL objective that allocates safety supervision along both axes. Adaptive prefix-tail weighting selects a response-specific prefix from cumulative divergence mass and gradually decays, rather than discards, tail supervision. English-anchored deficit calibration upweights target-language instances whose prefix-weighted divergence exceeds that of their paired English anchor. Both weights use predictions from a frozen, safety-conditioned copy of the student model, requiring neither an additional model nor a separate training stage. Across Qwen3 scales on MultiJail and PKU-SafeRLHF, DASD achieves the lowest average ASR among baselines for scale-benchmark pairs, while maintaining general capability on MMMLU and MGSM. Our code is available at https://anonymous.4open.science/r/12170F476.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.