acceptodds
Under review as a conference paper at ICLR 2027

Uncertainty as Heterogeneous Evidence: MUSE for Vision-Language Model Failure Prediction

Abstract

Vision–Language Models (VLMs) are increasingly used in high-stakes settings such as disaster re- sponse, where identifying unreliable predictions is essential for effective decision-making. Most current VLM failure-prediction methods rely on a single uncertainty perspective, which may over- look important reliability signals. We propose that VLM failures exhibit diverse reliability signatures that cannot be fully captured by one perspective alone. To address this, we present MUSE, a Multi- perspective Uncertainty Synthesis Ensemble that integrates five explicitly defined reliability perspec- tives: cross-modal alignment, semantic competition, intra-class geometry, class-conditional density, and neighborhood consistency. MUSE operates on frozen VLM representations without retraining or modifying the underlying model and does not require access to the training objective, enabling retrospective reliability estimation across pretrained backbones. Across 20 classification dataset– task settings, MUSE achieves mean AUROCs of 94.26%, 95.02%, and 94.36% on CLIP, SigLIP, and OpenCLIP, respectively, outperforming the strongest baseline by +7.99, +5.33, and +6.51 AU- ROC percentage points. Correlation analyses demonstrate distinct complementarity among these reliability perspectives, while cross-modal retrieval experiments further demonstrate applicability to real-world multimodal disaster-response scenarios. These findings support multi-perspective uncer- tainty synthesis as a robust approach to VLM failure prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.