The Multilingual Alignment Paradox: Preference Conflict Hurts Alignment but Prevents Language Collapse
Abstract
Preserving language-specific preferences is challenging under language-imbalanced training, especially when preferences conflict across languages. We study this interaction using a controlled preference dataset derived from global opinion data, varying language imbalance and cross-language preference conflict. Across four multilingual LLMs, we compare supervised fine-tuning (SFT) and direct preference optimization (DPO). As imbalance increases, SFT largely preserves the input language but loses preference accuracy, whereas DPO increasingly produces English responses to minority-language inputs, especially under shared preferences. This reveals a **multilingual alignment paradox**: preference conflict reduces preference accuracy but can protect language fidelity. At an English-to-each-minority ratio of 100:1, conflict reduces English fallback across all four models and raises mean joint language-and-preference accuracy after DPO from 43.2% to 62.1%. Objective analysis shows that, unlike SFT, within-language DPO comparisons leave output-language probabilities unconstrained at the distribution level, while training dynamics indicate stronger residual minority preference signals under conflict. Sparse autoencoder analysis further identifies stronger target-language and weaker English-associated activity under conflict. Intervening on these features reduces English fallback with limited impact on preference accuracy. Together, these findings suggest that language-selection features retained with language-specific preferences even under severe imbalance in the alignment process.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.