When Confidence Doesn’t Travel: Calibration of Open-Weight LLMs on African-Language MCQs Under Prompt Perturbations
Abstract
Research question: Are open-weight LLMs well-calibrated on African-language multiple-choice questions, does their calibration survive prompt perturbations, and at what compute cost? Answer: Mostly no. The failure is model-specific in ways aggregate benchmarks hide, and the full audit costs under 40 USD. We audit four open-weight instruct models (Llama-3.1-8B, Qwen2.5-7B, Gemma-2-9B, Mistral-7B) on 422,000 MCQ items spanning 24 Global-MMLU languages (9 African, 15 European) and 18 afrimmlu configurations, under 5 prompt perturbations, for a total cost of 39.48 USD on serverless A100s. Accuracy drops 16-31 points from European to African language cells for every model. The calibration gap, measured as Expected Calibration Error (ECE), is real but model-dependent at baseline: large for Qwen (+0.140 ECE [0.099, 0.177]) and Gemma (+0.133 [0.109, 0.159]), absent for Llama, and reversed for Mistral. Perturbation effects flip sign across models; a format that calibrates one model can break another, and Gemma loses 20 accuracy points in African language cells simply when asked to "answer with a single letter", with a 26.1% fallback rate across evaluated African cells that reaches 100% in some format_A afrimmlu cells. Miscalibration is largely repairable in-distribution: regularized logistic regression fit on approximately 250 items per cell cuts held-out ECE to 0.04-0.10 depending on region. We provide code, parsed outputs, and an executable reproduction notebook.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.