Multilingual Fine-Tuning via Localized Gradient Conflict Resolution
Abstract
Cross-lingual versatility has become a defining capability of modern Large Language Models (LLMs). However, fine-tuning these models often induces *negative interference*, where gains in one language come at another's expense. To address this, we cast multilingual fine-tuning as a multi-objective optimization (MOO) problem. Specifically, we introduce **Bucket-Level MOO**, a scalable distributed framework that applies gradient-based MOO algorithms locally, within the parameter buckets of distributed training. This enables conflict-aware updates without ever materializing a full gradient vector for any objective, and improves multilingual performance over standard fine-tuning by up to 4.3 points. Theoretically, we define *bucketwise Pareto stationarity*, show it is necessary for Pareto optimality under any fixed partition, and prove descent and convergence for representative variants. Our work brings standard MOO within reach for multilingual LLMs at a far lower cost. More broadly, Bucket-Level MOO is agnostic to both the objectives and the update rule, readily extending to other multi-task learning at scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.