AMIGO: Adaptive Multilingual Instruction Tuning via Gradient Similarity and Budgeted Optimization
Abstract
Multilingual instruction tuning requires allocating a limited training budget across languages whose cross-lingual interactions evolve throughout training. Prior work typically addresses data mixing (e.g., sampling ratios) and training strategy (e.g., curricula or scheduling) separately, often relying on static or heuristic decisions that do not account for the non-stationary nature of multilingual optimization. We propose **AMIGO** (Adaptive Multilingual Instruction Tuning via Gradient Similarity and Budgeted Optimization), a unified framework that couples language grouping, training order, and budget allocation through gradient geometry re-estimated as training progresses. AMIGO operates in regimes: at each stage, it clusters languages based on validation gradient similarity, selects the next language family via a compatibility score that balances optimization potential with transition smoothness, and allocates training budget using a continuous knapsack objective grounded in a first-order approximation of the multilingual loss. The utility of a language is defined by its need-weighted gradient alignment with out-of-family languages, while its cost is derived from validation loss and gradient norms, eliminating the need for external evaluation at regime boundaries. This formulation captures the dynamic nature of cross-lingual interactions and enables adaptive reallocation of training resources as optimization progresses. Empirical evaluations on various instruction-following sets across 23 languages and multiple model families demonstrate that _AMIGO_ outperforms existing data-mixing and data-selection baselines on average under fixed training budgets, with notable gains on several benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.