acceptodds
Under review as a conference paper at ICLR 2027

Verification, Not Correctness: Label-Free Distillation of Grounded Video Reasoning

Abstract

Video question answering (VideoQA) models often answer correctly while pointing to the wrong video evidence. Agentic reasoners close much of this grounding gap by verifying intermediate claims against observed events, at a steep inference cost. We introduce \method, which distills a verifier-equipped agentic teacher into a student that emits an answer, typed temporal evidence, and a self-verification score—calibrated to predict the verifier's acceptance of its output—in a single autoregressive pass. A controller trained by preference optimization on the verifier's signal alone raises the teacher's verified yield; a tilted verification filter selects and weights traces without answer labels and distills outputs and belief-state representations; and a self-verification head amortizes the verifier itself, enabling a cost-controlled escalation cascade. On NExT-GQA the 4B student reaches \accgqa, within of the teacher's , at its inference cost, and improves on correctness-filtered distillation by \accgqa without answer labels in controller training or trace selection; more than half of this gain comes from a higher localization rate among correct answers. The advantage replicates on a second teacher–verifier pair and on ReXTime and CG-Bench. A supporting analysis expresses the filtered teacher's correct-but-ungrounded mass through measurable verifier primitives and transfers the bound to the student; we report it as a population-level diagnostic with stated limitations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.