LARA: Language-Adaptive Reward-Anchored Preference Optimization
Abstract
Preference optimization is typically applied to multilingual language models with one global margin scale, implicitly assuming that chosen-rejected log-probability gaps are comparable across languages. This assumption is fragile since tokenizer fertility, script, response length, and language entropy make the same semantic preference induce different numerical margins. Existing remedies reshape the loss or reweight examples, yet they still apply one global margin scale to every language. We therefore introduce Language-Adaptive Reward-Anchored Preference Optimization (LARA), which instead changes the unit of optimization while preserving the implicit-reward view of direct preference optimization (DPO). Concretely, LARA scores completions with tokenizer-independent effective-length normalization, calibrates margins with per-language scales estimated from reference gaps, and optionally adds cross-lingual anchoring and language-level group distributionally robust optimization (group-DRO). Across three backbones (Qwen3-4B, Aya Expanse 8B, and Gemma-4-12B-it) and seven languages, LARA improves average and worst-language held-out accuracy over every baseline while also balancing languages. On Gemma, where DPO barely moves the base model (51.6%/47.0%), LARA reaches 68.0%/67.3% against 60.7%/56.5% for odds-ratio preference optimization (ORPO), with no length bias and the smallest language gap of any method. Under a Friedman-Nemenyi rank test on Gemma, LARA outranks every baseline at the 5% level. Ablations and theory then isolate the mechanism, identifying effective-length normalization as the dominant driver and per-language calibration as a smaller lower-tail and balance gain, while diagnostics rule out verbosity. The margin unit therefore decides how much accuracy a multilingual objective recovers, and the per-language scale decides how evenly that accuracy spreads across languages.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.