acceptodds
Under review as a conference paper at ICLR 2027

PROBABILISTIC EVALUATION OF LARGE LANGUAGE MODELS WITH AMORTIZED BAYESIAN RELIABILITY ESTIMATION

Abstract

Large language model (LLM) evaluation increasingly relies on automated performance metrics to compare models across tasks, yet the reliability and uncertainty of these evaluations remain insufficiently characterized. Conventional evaluation metrics typically provide point estimates of model performance, making it difficult to determine whether an observed performance difference reflects a meaningful improvement or uncertainty in the evaluation process. In this work, we propose a Bayesian inference-based performance metric for quantifying the reliability of LLM evaluation. The proposed approach models evaluation outcomes probabilistically, enabling performance estimates to be accompanied by uncertainty information and providing a principled basis for assessing the confidence and reliability of model comparisons. We evaluate the proposed metric across multiple language models and benchmarks, including PolyGemma, BERT, and Polyglot, and compare its behavior against established evaluation approaches. Our experimental analysis demonstrates how Bayesian inference can expose uncertainty that is otherwise hidden by conventional point-based metrics and provide a more informative characterization of model performance. The results highlight the importance of accounting for uncertainty when ranking or comparing LLMs and suggest that reliable evaluation should consider not only the magnitude of measured performance but also the statistical confidence associated with that measurement. The proposed framework provides a general and interpretable approach for uncertainty-aware evaluation of LLMs and can be applied to systematic comparisons across models, datasets, and evaluation settings.We formulate LLM evaluation as Bayesian inference over a continuous latent semantic-quality variable, using a Gaussian prior to represent prior beliefs and uncertainty over graded semantic relatedness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.