Decision-Aligned Evaluation of Uncertainty Quantification
Abstract
Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions. We introduce *decision-alignment*, a criterion that reveals which metrics faithfully rank models according to their decision-making performance. Applying this framework, we show that many widely used uncertainty evaluation metrics are either misaligned with common decision problems or encode implausible prior beliefs about the downstream task. We then propose *prior-weighted utility metrics*, a special class of proper scoring rules that provides decision-aligned uncertainty evaluation. Across benchmark experiments and real-world case studies, our metrics consistently align with realized decision utility, while conventional metrics do not. Our results surface gaps in the current UQ evaluation protocol and offer a principled evaluation scheme toward decision-relevant UQ evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.