LEAD: Models Know the Language, but Token Preferences Do Not Always Follow
Abstract
Multilingual large language models often experience *language confusion*, failing to consistently generate responses in the intended language. Existing mitigation approaches often require costly model retraining or inference-time interventions that assume access to target language labels. Moreover, they do not distinguish whether language confusion stems from a failure to encode the intended language or from a failure to utilize the encoded language information during generation. This paper introduces **Language Evidence-Aware Decoding** (LEAD), a lightweight method that exploits language-subspace logits from the final hidden state to selectively intervene only at decoding time. Our method is motivated by the v-info theory and three empirical findings: linear probes identify languages with high accuracy, the logit contributions from language identity are weak, and causal interventions on the LM head reduce language confusion. Experiments on multiple multilingual LLMs across diverse languages demonstrate that LEAD substantially reduces language confusion while preserving general task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.