APL-OPSD: Instance-Adaptive Privileged Language Selection for Multilingual On-Policy Self-Distillation
Abstract
Large language models acquire knowledge from multilingual corpora, but this knowledge is not equally accessible through every language. The most informative language can vary across questions, domains, and cultural contexts, making fixed English supervision suboptimal. We propose Adaptive Privileged Language On Policy Self Distillation (APL-OPSD), a framework that selects the language of privileged context separately for each training instance. Given parallel question and answer pairs in multiple languages, APL-OPSD performs a one time offline probe that measures how each candidate context changes the likelihood of the target language reference answer and selects the most informative non target language. During training, the student observes only the target language question and generates an on policy response, while a frozen teacher additionally observes the selected crosslingual context and provides dense token level supervision over the same trajectory. Experiments across multiple languages, model scales, and tasks show that APL-OPSD achieves the strongest aggregate performance among the compared methods. Further ablations suggest that informative language selection contributes beyond merely adding privileged context. APL-OPSD jointly trains a single multilingual model and requires no additional routing, translation, or reference information at inference time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.