acceptodds
Under review as a conference paper at ICLR 2027

Multilingual Fine-Tuning via Localized Gradient Conflict Resolution

Abstract

Cross-lingual versatility has become a defining capability of modern Large Language Models (LLMs). However, fine-tuning these models often induces *negative interference*, where gains in one language come at another's expense. To address this, we cast multilingual fine-tuning as a multi-objective optimization (MOO) problem. Specifically, we introduce **Bucket-Level MOO**, a scalable distributed framework that applies gradient-based MOO algorithms locally, within the parameter buckets of distributed training. This enables conflict-aware updates without ever materializing a full gradient vector for any objective, and improves multilingual performance over standard fine-tuning by up to 4.3 points. Theoretically, we define *bucketwise Pareto stationarity*, show it is necessary for Pareto optimality under any fixed partition, and prove descent and convergence for representative variants. Our work brings standard MOO within reach for multilingual LLMs at a far lower cost. More broadly, Bucket-Level MOO is agnostic to both the objectives and the update rule, readily extending to other multi-task learning at scale.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.