acceptodds
Under review as a conference paper at ICLR 2027

SPAR: LEARNING VLM RISK FROM STRUCTURED PEER AGREEMENT

Abstract

Vision–language models can be confidently wrong, and their confidence scales transfer poorly across architectures and tasks. We study fixed-focal selective prediction, where heterogeneous peer outputs help decide whether to accept or abstain on a focal VLM answer without changing its content. Common multi-VLM uncertainty scores compress this evidence into scalar agreement, even though the same mean disagreement can arise when peers agree with one another but oppose the focal answer, or when they disagree among themselves. We introduce SPAR, a Structured Peer-Agreement-based Risk estimator for fixed-focal selective prediction. It decomposes focal disagreement into focal isolation and peer ambiguity, and combines this decomposition with relative peer confidence. Percentile-rank alignment makes continuous features comparable across models and datasets, allowing a single risk scorer to be applied unchanged across different settings. Across five focal VLMs and three multimodal benchmarks, SPAR achieves the best aggregate AURC and consistently outperforms the evaluated peer-based and repeated-sampling alternatives. Under simultaneous model-family and dataset shift, it reduces macro-AURC by 7.05% over the strongest baseline (Consensus Entropy), with gains in every setting, while achieving up to a 10.1× generation-time speedup over the strongest repeated-sampling comparison. These results establish structured output relations as an effective and transferable reliability signal for selective VLM prediction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.