Discrimination Is Not Selection: Rethinking Verifier Evaluation for Multimodal Test-Time Scaling
Abstract
A verifier can rank correct responses above incorrect ones on average and still select the wrong answer. This distinction matters for multimodal test-time scaling, where the value of additional sampling depends on recovering the correct answers it produces. We investigate this verification–selection gap and identify a mismatch between how verifiers are evaluated and how their scores are used. Pooled discrimination compares candidates across queries that never compete for selection, while within-query ranking averages over pairs even though the final decision depends on the score maximum. We formalize these distinctions and derive an exact decomposition of selection loss into incorrect candidates that strictly dominate and mixed-quality ties that are resolved incorrectly. Across multimodal reasoning benchmarks, informative verifier scores coexist with substantial failures to recover available correct answers, with different failure patterns across scoring interfaces. Broader candidate sampling and targeted interventions over frozen representations do not consistently eliminate this loss. Our findings motivate treating the recovery of available correct answers as an explicit evaluation target: verifiers for test-time scaling should be assessed through the decisions their scores induce, alongside their ability to discriminate correctness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.