Stop Early, Spend Less: Equal-Accuracy Frame-Budget Saving via Agreement-Gated Adaptive Sampling in Video Language Models
Abstract
Video large language models (VLMs) answer questions over long videos by sampling a fixed number of frames. Practitioners pick this frame budget a priori, yet the accuracy-vs-frames curve saturates at a domain-dependent point: beyond it, extra frames yield no accuracy but multiply compute, and can even hurt. We show that a simple, training-free agreement-gated rule—answer with a few frames, and if a slightly larger sample agrees, stop early; otherwise escalate—matches the accuracy of a strong static-uniform baseline at its best global budget while spending far fewer frames. On Video-MME with five VLMs spanning three families and three scales (Qwen2.5-VL-7B; Qwen3-VL-8B; InternVL3-8B; InternVL3.5-8B; InternVL3.5-4B), STOP-EARLY is statistically non-inferior to static sampling at 24 frames (TOST, margin ±3 percentage points) while using approximately 41–48% fewer frames on three of the five models (the exception is the model with the steepest short-video curve). Across five video-QA settings (three Video-MME splits, TempCompass, MVBench), the benefit is a monotone function of the static curve's slope: the strong form (near-perfect rank correlation, Spearman rho at most -0.9) holds on three of the five models, and the weaker sign-level form (a negative slope predicts a free lunch) holds on all five. Where the curve is flat or declining, STOP-EARLY saves 45–54% of frames at equal or better accuracy (on TempCompass, +4.0 percentage points at -54% frames, and up to +5.5 percentage points on other models); where it rises steeply, the fixed high budget is genuinely needed—a predictable scope boundary. Since the optimal fixed budget drifts across domains and is unknown at deployment, a fixed budget tuned on one domain incurs regret on another, while STOP-EARLY tracks the per-domain optimum online. A rate-matched random gate fails non-inferiority, isolating that the gain comes from knowing when to stop. We release all code, per-item results, and analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.