What English-Only, Non-Frontier Evaluation Misses About Consistency-Based Confidence Estimation
Abstract
Confidence estimation helps users decide when to trust an AI model's answer. Most common black-box confidence estimation methods, such as self-consistency, measure agreement among answers sampled from the model given the question. Analyzing recent work, we find that research results on consistency-based methods largely come from English-only evaluations of small models, leaving other languages and recent models under-explored. We therefore evaluate consistency-based confidence estimation across nine languages and a range of small open-source to frontier models. Our results show that self-consistency performs worse outside English for all 17 models we tested. Moreover, verbalized confidence using recent Claude models outperforms self-consistency, despite using only one LLM call. For a more effective consistency-based approach across models and languages, we introduce Rosetta, which uses translation as a principled perturbation method and measures agreement between the resulting multilingual samples and the original answer to estimate confidence. Rosetta improves confidence estimation over same-language rewording in English and outside English at a matched call budget and without labeled calibration data, while narrowing the average language gap.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.