acceptodds
Under review as a conference paper at ICLR 2027

PolyAEGIS: Benchmarking Large Language Model Safety Across 100 Languages

Abstract

Large language models are increasingly deployed across linguistically diverse populations, yet safety evaluation remains concentrated in English and a small set of high-resource languages. We introduce a multilingual safety benchmark derived from the AEGIS unsafe-prompt taxonomy, translated into 100 typologically diverse languages spanning 16 language families, 26 scripts, and a balanced split of high- and low-resource languages. Source prompts are filtered from the AEGIS corpus across 21 safety-hazard categories, machine-translated using NLLB-200, and scored for translation quality using three complementary metrics (BERTScore-F1, COMET-Kiwi, chrF++), with a translation retained unless it simultaneously fails all three. The benchmark is constructed via multi-factor equal-allocation stratified sampling across safety category and prompt-length strata, crossed completely with all 100 target languages, yielding full coverage — 100/100 languages, 21/21 categories, 16/16 language families, and 26/26 scripts — across more than 9,000 prompt-translation pairs. We generate model responses using three open-source large language models spanning distinct architectures and parameter scales, and evaluate safety compliance using an LLM-based judging protocol combining closed- and open-source judges. We present the dataset construction methodology, translation-quality validation, and coverage analysis here, with full response-generation and safety-evaluation results to follow in the complete submission.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.