Cross-lingual Multi-view On-Policy Self-Distillation:Many Voices, One Model, No New Knowledge
Abstract
Multilingual language models often exhibit weaker reasoning capabilities in low- resource languages. Cross-lingual reasoning can leverage high-resource-language renderings of the same problem at inference time, but doing so requires repeated multilingual decoding and answer aggregation for every query. We ask whether this inference-time benefit can instead be internalized into a model that reasons di- rectly in the target language. We introduce Cross-lingual Multi-view On-Policy Self-Distillation (CMOPD), which uses multiple language versions as privileged training-time views of the same problem. The views provide dense token-level feedback on one low-resource on-policy trajectory, and on-policy distillation con- solidates this feedback into a single target-language policy. After training, the model answers directly from the low-resource query without multilingual decod- ing or answer aggregation. Across Qwen3 1.7B, 4B, 8B backbones and four AfriMGSM target languages (Swahili, Igbo, Yoruba, Amharic), CMOPD with English and Chinese views improves mean pass@12 over our reproduced single- view baseline at every scale (1.7B: +0.2, 4B: +2.6, 8B: +1.8), while per-language effects remain heterogeneous. On Qwen3-8B Swahili, adding the Chinese view yields similar gains during training (+4.8) and test-time synthesis (+4.4).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.