acceptodds
Under review as a conference paper at ICLR 2027

WHEN VIDEO VERIFIER SCORES MISLEAD: CAPABILITY ASSESSMENT AND VERIFIER SELECTION

Abstract

We audit video verifiers, models that judge whether a candidate answer or trajectory agrees with a video, and show that their scores can mislead both conclusions about model ability and verifier selection. We vary two evaluation choices—negative type and judgment interface—while holding the tested models fixed. On controlled tasks, direct judgment yields win rates of 0.78–0.94 against visible-start negatives but only 0.46–0.58 against endpoint-matched negatives that require checking the internal trajectory. Yet on paired videos where the correct candidate reverses with the internal evidence, CoT judgment raises paired success from 0–6% to 75–88% for three of four models, showing that poor direct-judgment performance need not imply an absence of evidence-responsive discrimination. This CoT gain does not transfer to natural video: across 400 videos, CoT judgment yields no significant accuracy gain for any model and significantly reduces accuracy against event-denial negatives for three of four. Negative type also affects verifier selection: when content-matched negatives are the target, selecting on event-denial performance reduces held-out accuracy by 8.9 percentage points. These results separate success on the negative type tested, ability that can be elicited from the model, and suitability for the intended verification target.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.