DECAF: Disentangling Capability and Translation Artifacts for Multilingual LLM Evaluation
Abstract
While most large language models (LLMs) support a variety of languages beyond English, accurately evaluating their language-specific capabilities remains a challenge. Multilingual benchmarks are central to this evaluation; however, because they are typically translated from an English baseline using machine translation, the translation process often introduces item-level artifacts, such as translation errors and culturally sensitive content. These translation artifacts obscure the accurate estimation of a model's language-specific capabilities. To address this challenge, this paper introduces the Multilingual Differential Item Functioning Item Response Theory (MDIF-IRT) model. The MDIF-IRT framework advances multilingual LLM evaluation by disentangling language effects into distinct model capability shifts and item difficulty shifts. An application to the MMLU-ProX dataset, evaluating 25 models across 29 languages, demonstrates the utility of the proposed method. The MDIF-IRT model estimates language-specific latent capabilities while accounting for item-specific cross-lingual difficulty variation to improve estimation of capability and item parameters. Furthermore, the framework provides a tool for identifying benchmark items with significant cross-lingual translation artifacts and improves prediction of unobserved item responses compared to empirical accuracy and conventional IRT models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.