NuosuBench: Revealing the Limits of Multilingual Large Language Models in Standardized Yi
Abstract
Multilingual large language models (LLMs) exhibit strong cross-lingual transfer across many languages, but it remains unclear whether this capability extends to languages written in severely underrepresented scripts. We investigate this question through Standardized Yi (Nuosu), a language spoken by millions of people in southwestern China but largely absent from existing LLM training and evaluation ecosystems. We introduce Nuosu-Datasets and Nuosu-Benchmark, the first large-scale resource suite for training and systematically evaluating Standardized Yi language models. Nuosu-Datasets contains 78.6M tokens for continual pre-training and 166M tokens comprising 1.1M instruction instances for supervised fine-tuning and cross-lingual alignment. The data are collected from authoritative web and expert-provided sources through OCR-based digitization, multi-stage filtering, and expert validation. Nuosu-Benchmark contains 110513 instances across five sub-benchmarks that evaluate lexical knowledge, sentence-level understanding, culturally grounded knowledge, administrative language, and educational comprehension using 12 complementary metrics. Evaluating 14 representative LLMs reveals a substantial gap between general multilingual capability and competence in Standardized Yi. All evaluated general-purpose models achieve aggregate scores below 40%, indicating that model scale and broad multilingual training alone do not ensure effective transfer to an underrepresented script. In contrast, BIMO-8B, an 8B model adapted using Nuosu-Datasets, achieves 50.1%, outperforming the strongest evaluated general-purpose model by 10.2 percentage points. Further analyses show that language-specific training produces improvements across multiple levels of linguistic competence, while culturally grounded and knowledge-intensive tasks remain particularly challenging. Our resources and findings establish a reproducible foundation for studying cross-lingual transfer, data-efficient adaptation, and evaluation in extremely low-resource language modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.