acceptodds
Under review as a conference paper at ICLR 2027

OMNILINGUALGAIA2: EVALUATING THE MULTILINGUAL GAP IN FRONTIER AI AGENTS

Abstract

Agentic benchmarks aim to measure how well AI agents plan, reason, and act in realistic environments, yet they are almost exclusively English-only. As AI agents are increasingly deployed to linguistically diverse users, whether agentic competence measured in English generalizes to other languages remains an open question. We introduce OMNILINGUALGAIA2, a machine-translated expansion of the GAIA2 agentic benchmark, with partial human-expert validation. OMNILINGUALGAIA2 covers ten target languages spanning five writing systems and includes a cross-lingually calibrated multilingual verifier. Evaluating seven frontier proprietary and open-weight models, we find a universal cross-lingual gap of 8.8–18.4 pass@3 points. The gap varies substantially across agents, is concentrated in tool orchestration rather than quantitative reasoning, and persists with increasing model scale. A stratified error attribution shows that at least 55% of the observed gap can be attributed to model-driven failures, which we further study through human-expert linguistic analysis. We identify morphological cue loss and amplified ambiguity as the primary failure mechanisms, especially in non-Latin-script languages. These findings demonstrate that strong English-language agentic performance does not reliably transfer across languages, and argue that multilingual agentic evaluation should become a standard component of reporting for globally deployed agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.