ProbabilityBench: Model Evaluation and Calibration with Exact Ground-Truth Uncertainty
Abstract
Claims about predictive uncertainty are hard to check against real data: each example typically carries one realized label, never the distribution behind it, so uncertainty quantification is often judged against estimates rather than truth. We introduce ProbabilityBench with three simulated worlds, i.e., credit review, peer review, and covert transfer, where partial observability creates irreducible uncertainty, natural-language rendering provably preserves it, and every case's exact class probability is known. Such ground truth has existed only in toy settings such as Gaussian class conditionals, unsuited to testing a language model. With ProbabilityBench, a model's probability is scored by its exact KL divergence, e.g., Jev is both the cheapest and the closest to the true probability in every world among the models we compared. Ground-truth probabilities also enable us to evaluate the evaluation methods themselves, to assess how accurate uncertainty-quantification methods such as calibration estimators and Bayes-error estimators actually are.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.