acceptodds
Under review as a conference paper at ICLR 2027

Bug or Feature? Cross-Lingual Preference Distillation, Optimization Dynamics, and Capacity Contention in Multilingual LLMs

Abstract

Multilingual Large Language Models project diverse languages into shared representations to transfer task capabilities, yet standard post-training alignment predominantly leverages English-centric preference data, imposing Western-centric defaults or triggering language confusion and English drift on non-Western queries. We hypothesize that an unaligned model's own greedy continuation modes provide effective negative supervision to purge degenerative artifacts and default priors without requiring additional human preference curation. To test this, we evaluate Cross-Lingual Preference Routing and Self-Contrastive Mode Suppression (CP-SMS), a post-training framework that pairs localized native completions () with the base model's unaligned greedy continuation modes () alongside general capability data. To benchmark performance without proprietary judge drift, we formalize a triangulated protocol anchored by multi-task language competence (Belebele), an open-weights Contrastive Perplexity Probability probe (CPP, ) measuring relative conditional likelihood under localized versus Western conditioning prefixes, and zero-shot frontier evaluation. Across five foundational model families and typologically diverse languages, our experiments reveal that supervised fine-tuning acts solely as a positive syntax booster, whereas self-contrastive preference tuning repels unaligned default modes, with sequential alignment (SFT_to_DPO) achieving the most favorable empirical balance between reading comprehension and preference win rates. Crucially, we uncover that in low-resource languages, DPO acts primarily as a generation-stabilization operator, purging repetitive loops, prompt echoing, and English drift, and in high-resource languages, it steers pragmatic completion stance away from English-dominant pre-training defaults. Furthermore, mixture scaling demonstrates that preference tuning acts as a latent steering operator rather than synthesizing absent representations as low-resource languages collapse into Negative Preference Transfer due to the compounding interaction of pre-training token deficits and extreme subword fragmentation. Finally, there is a significant trade-off between the subspace rank and the number of languages in post-training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.