Confidence Needs Evidence: Reliable Confidence Estimation for Multimodal Models
Abstract
Reliable deployment of multimodal models requires confidence estimates that distinguish trustworthy predictions from those that should be rejected. However, false confidence—the assignment of high confidence to incorrect predictions—remains difficult to mitigate. Our experiments show that many false-confidence cases in multimodal question answering arise from language priors without corresponding task evidence, indicating that confidence is unreliable when it lacks evidence support. Motivated by this finding, we propose , a simple, training-free confidence estimator. With only two forward passes, measures evidence support through the full-versus-blind probability lift of the fixed full-input prediction and combines this signal with normalized predictive-entropy confidence to produce an evidence-weighted confidence score. We evaluate with multiple models on four benchmarks covering video QA and joint audio–video QA. consistently outperforms state-of-the-art baselines and achieves the highest correctness AUROC in all 16 benchmark–model settings, demonstrating its reliability and generality for multimodal confidence estimation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.