acceptodds
Under review as a conference paper at ICLR 2027

Calibrated Answer-Stability Halting for Active Long-Video Understanding

Abstract

Long-video question answering enables efficient access to lengthy recordings, but confidence from partial observations can be misleading when later evidence changes the answer. We seek to reduce video processing while controlling premature commitment to answers that would change under a fixed, expanded view. We introduce Stability-Based Halting (SBH), which uses a lightweight linear predictor combining current confidence, observation progress, and coarse cues from remaining video content to estimate the current answer's advantage over alternatives under that reference view. Calibration across videos yields a lower bound that triggers stopping when positive and controls early-stop reference disagreement across steps and questions under exchangeable video sampling. On held-out Video-MME videos with Qwen2.5-VL-7B, SBH achieves a speedup over scoring all 64 budgeted frames once, with 68.02% versus 67.65% accuracy and 5.19% observed video-level early-stop reference disagreement at a 10% target.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.