Shared Computation Is Not Shared Access: Diagnosing and Improving Multilingual Reasoning
Abstract
Multilingual models can solve a problem in one language yet fail on the same problem in another. Final-answer accuracy leaves unclear whether the difficulty lies in executing calculations or in identifying the required quantities and operations and using intermediate results. To distinguish these difficulties, we develop a controlled framework that varies cross-language correspondence and the information supplied to the model. We measure autonomous solving from the question, conditional computation of a supplied operation, and continued solving from a correct partial solution. Experiments with synthetic languages, English–Chinese tasks, and two open-source model families show that accurate local computation can coexist with unreliable autonomous solving. Further diagnostics reveal difficulties in retrieving numerical inputs and completing solutions beyond a correct start. Guided by these diagnostics, we train Qwen3.5-4B first on ordered arithmetic steps extracted from natural-language solutions and then on full target-language chain-of-thought. Compared with chain-of-thought training in both stages, this route achieves higher mean autonomous accuracy under matched source exposure and training updates. These findings motivate evaluating and training multilingual models for both accurate computation and the ability to build complete solutions from a question.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.