L-MoL: Mitigating Multilingual Interference via a Language-Guided Mixture of Low-Rank Experts
Abstract
Can a trained dense multilingual translation model benefit from sparse adaptation without acquiring new training data? We propose L-MoL, a language-guided mixture of low-rank residual experts attached to decoder feed-forward networks. The adaptation protocol freezes the dense backbone and trains the expert branches and router. Concatenating a target-language embedding with each token representation gives the router a language-specific logit bias while retaining token-dependent selection. On the 188 available OPUS-100 test directions, L-MoL improves its dense seeds by 2.69 BLEU (Base) and 1.54 BLEU (Large), with additional parameters equivalent to 3.8% and 8.3% of the respective seed sizes. At Base scale, it exceeds the strongest tested language-agnostic routing adaptation by 1.84 BLEU and its own hidden-state-only ablation by 1.95 BLEU. On five selected language pairs, mean routing-overlap reductions are approximately 81% for the three distant pairs and 7.5% for the two related pairs. These results support explicit target-language conditioning as a useful inductive bias for multilingual expert allocation; routing overlap is a diagnostic, not a direct measurement of gradient conflict.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.