Visual Contrast for Reliable Uncertainty in Label-Free VLM Ranking
Abstract
Selecting a vision-language model (VLM) for a downstream task is challenging when ground-truth labels are unavailable. Existing uncertainty-based ranking methods use signals such as negative log-likelihood (NLL) and entropy, but overconfident predictions can make these signals unreliable. We observe that excessive reliance on either linguistic priors or visual information is associated with greater overconfidence than reliance on both sources. Motivated by this observation, we propose Visual Contrast for Reliable Uncertainty (VCRU), a label-free method for ranking VLMs on multiple-choice question-answering tasks. VCRU compares the NLL of the original first-token prediction and the output entropy between original and visually degraded inputs. It identifies confidence conflicts, where NLL and entropy shift in opposite directions, as diagnostic signals of potentially unreliable confidence. A drift gate then emphasizes predictions with high original confidence and substantial changes in relative confidence, yielding a model-level penalty that adjusts the baseline ranking score. Experiments with 36 heterogeneous VLMs across four benchmarks show that VCRU consistently improves upon its underlying uncertainty baselines and achieves the strongest ranking correlations on three benchmarks. On MMStar, it improves Spearman's correlation from 0.8273 to 0.9437 over first-token NLL. These results demonstrate the value of visual contrast for assessing the reliability of confidence in label-free VLM ranking.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.