MACE: Affinity and Conflict-Aware Data Selection for Multilingual Instruction Tuning
Abstract
Multilingual instruction tuning improves the ability of large language models to follow instructions across languages. The training mixture shapes both the supervision available to each language and interactions between languages. Data selection is therefore central to effective multilingual training. Through shared parameters, however, an example aligned with learning directions in its own language may induce updates that oppose those of other languages. Selection requires reliable references for diverse within-language directions and an explicit assessment of cross-language opposition. We introduce MACE (Multilingual Affinity and Conflict-aware Example Selection), a gradient-based data selection framework addressing both requirements. MACE constructs multiple gradient prototypes for each language and measures a candidate's in-language affinity through alignment with its closest prototype. It assesses cross-language conflict exposure using the Magnitude-weighted Gradient Conflict Ratio (MGCR), which measures the share of gradient-product magnitude contributed by opposing coordinates. Combining affinity with an exposure penalty yields a score for ranking and retaining candidates within each language. Across five multilingual benchmarks, MACE + KMC improves the average score over LangGPS + KMC by a relative 1.99% on Qwen2.5-7B and 4.59% on Llama-3.1-8B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.