Language Shapes Decisions in Large Language Models
Abstract
Large language models are deployed globally as multilingual decision-makers under the assumption that language is a neutral doorway to the same underlying system. Although leading models often reach similar factual accuracy across languages, it remains unclear whether they make the same choices when the language changes. We tested this question with 52,800 controlled forced-choice trials covering six model families, eleven scenarios, and seven behavioral domains. Each model was compared with itself in matched English and Chinese conditions, and we also asked it to respond in a language different from the language of the scenario. Across 66 comparisons, 37 showed a statistically significant shift after Benjamini–Hochberg correction at , or 56.1% of all comparisons. The arcsine-transformed contrast reached . The direction of change depended on the model and sometimes reversed completely. In a loss-framed risk task, one model moved from complete risk tolerance in English to complete risk aversion in Chinese, while another model showed the opposite pattern. Agent attribution was the only domain with the same directional shift across all six models. When the requested answer language was crossed with the scenario language, mixed-language responses frequently departed from the monolingual envelope by more than , with marked suppression in some pairs and marked amplification in others. These results show that the distribution of observable decisions can change with language and support evaluating model behavior separately by language before deployment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.