acceptodds
Under review as a conference paper at ICLR 2027

ANCHOR: Cross-Lingual Reasoning with Anchored Reinforcement Learning Under Accuracy-Language Trade-off

Abstract

Chain-of-Thought LLMs support multilingual reasoning, yet on low-resource language questions, they often rely on high-resource languages such as English for accuracy. Forcing reasoning in the question language improves readability for question-language speakers but may compromise accuracy. To address this trade-off, we propose Anchor, a representation anchoring and behavioral framework for cross-lingual reasoning. We structure the response into three segments: a multilingual reasoning segment, a question-language public explanation, and a verifiable answer. At the representational level, layer-wise probing reveals a recovery pattern in question language signals, based on which we introduce a one-sided anchoring loss to preserve question language routing; at the behavioral level, we design a multi-dimensional reward with two language-gated rewards to shape language choices under answer correctness and format validity. Both levels are jointly optimized in a DAPO-based policy optimization process. Experiments show Anchor enables flexible cross-lingual reasoning on multiple benchmarks. On MGSM, Anchor reaches 86.07% average accuracy across 11 languages, delivering 1.90% and 1.94% gains over the baseline for 4 low-resource and 7 high-resource languages respectively, alongside 74.2% public-segment question-language alignment. On MMLU-ProX-Lite and Global-MMLU-Lite, Anchor boosts performance on 41 of 52 languages and modest alignment improvements.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.