acceptodds
Under review as a conference paper at ICLR 2027

Free Generation-Time Signals Gate a Video LLM Better Than a 32B Judge

Abstract

Video large language models sample many candidate answers, and a large share of those candidates are wrong, so a deployed system has to decide for each answer whether to act on it. We ask which verification signal should make that decision. Two signals the sampling protocol already produces, the model's own confidence and its self-consistency across samples, separate right answers from wrong ones better than a 32B vision-language judge run once per question: they beat the judge in 6 of 8 model-dataset cells with question-level bootstrap intervals excluding zero, at no inference cost beyond the samples already drawn. The evidence is a scored pool of 36,000 labelled candidates from two open backbones on four benchmarks, run through five families of verification signal (confidence, self-consistency, the judge, frame-evidence agreement, and temporal metamorphic probes), each with a finite-sample false-accept gate and its inference cost recorded. The two video-specific families are the weakest, and combining all five adds little over the best single signal. Gating on any of these signals needs one repair: on tied multiple-choice scores a naive split-conformal gate exceeds its false-accept bound in nearly every single-signal setting, and standard randomized tie-breaking restores validity within seed noise. We release the scored pool so that a new signal or gate can be compared against the free baselines without a GPU.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.