acceptodds
Under review as a conference paper at ICLR 2027

The Multilingual Alignment Paradox: Preference Conflict Hurts Alignment but Prevents Language Collapse

Abstract

Preserving language-specific preferences is challenging under language-imbalanced training, especially when preferences conflict across languages. We study this interaction using a controlled preference dataset derived from global opinion data, varying language imbalance and cross-language preference conflict. Across four multilingual LLMs, we compare supervised fine-tuning (SFT) and direct preference optimization (DPO). As imbalance increases, SFT largely preserves the input language but loses preference accuracy, whereas DPO increasingly produces English responses to minority-language inputs, especially under shared preferences. This reveals a **multilingual alignment paradox**: preference conflict reduces preference accuracy but can protect language fidelity. At an English-to-each-minority ratio of 100:1, conflict reduces English fallback across all four models and raises mean joint language-and-preference accuracy after DPO from 43.2% to 62.1%. Objective analysis shows that, unlike SFT, within-language DPO comparisons leave output-language probabilities unconstrained at the distribution level, while training dynamics indicate stronger residual minority preference signals under conflict. Sparse autoencoder analysis further identifies stronger target-language and weaker English-associated activity under conflict. Intervening on these features reduces English fallback with limited impact on preference accuracy. Together, these findings suggest that language-selection features retained with language-specific preferences even under severe imbalance in the alignment process.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.