Does High Accuracy Imply Trustworthiness? A Large-Scale Evaluation of Confidence and Risk in Frozen Vision and Video Foundation Models
Abstract
As large-scale vision and video foundation models enter high-risk settings, a standard safeguard is selective prediction, where a model acts only on high-confidence outputs while deferring uncertain ones. However, accuracy alone does not reveal whether a model’s most confident predictions are reliably correct. Common uncertainty metrics also do not directly address this issue. AUROC and expected calibration error (ECE) evaluate confidence across the full prediction set, FPR@95 evaluates at a fixed error-rejection target, and other metrics remain confounded by model accuracy. In this work, we focus on the trustworthiness of a model's confidence levels among predictions selected for action. We propose that the positive likelihood ratio (LR+), originally defined for diagnostic testing, can be adapted to measure how much more often correct than incorrect predictions are included among a fixed fraction of the most-confident predictions. We evaluate LR+ across 18 frozen vision and video foundation-model backbones from six families and three scales, spanning image classification, semantic segmentation, video classification, and action anticipation. On ImageNet-1K, DINOv3 ViT-H+ and SigLIP2 g-opt achieve nearly identical top-1 accuracy (88.4% vs. 88.1%), yet their LR+ values among the top 20% most-confident predictions differ by more than twofold (24.0 vs. 49.7). Across all 18 models, AUROC and FPR@95 show little variation and rank models differently from LR+, which reveals substantial differences. Across downstream tasks, the relationship between accuracy and LR+ changes, and different models achieve the highest LR+. These results show that the trustworthiness of a model based on its confidence scores cannot be inferred from accuracy alone or assumed to transfer across benchmarks, and should be evaluated for each task at the fraction of predictions intended for action.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.