acceptodds
Under review as a conference paper at ICLR 2027

Video Understanding at a Glance: What Remains at Two Tokens per Frame?

Abstract

Embodied and wearable systems often need a fast, low-cost understanding of ongoing visual experience before finer-grained perception or action is required. We formulate this need as *glance video understanding*: forming an overall understanding of scene context, salient events, and their temporal evolution—rather than exhaustively recovering every visual detail—at low visual cost. Preserving every frame as a distinct unit of visual evidence, we ask whether such understanding can still be sustained when the available visual-token budget is pushed to an extreme. Unlike an isolated image, a video frame carries both its *current visual state* and its *temporal contribution* relative to preceding observations. Compressing both into a single token would force these two aspects to compete within the same representational slot; we therefore set the minimum per-frame budget to two tokens—one for *what is present* and one for *what is changing*. This leads to our central question: *what remains of video understanding under this purely visual, query-independent boundary?* We investigate this question with **GlanceVLM**, which pairs an **Anchor Token** for global semantic context with an **Innovation Token** for temporally distinctive evidence. Difference-guided innovation encoding uses inter-frame changes to select informative regions while retaining their current-frame visual content, allowing both tokens to be formed before any downstream question is known. Across eight benchmarks, GlanceVLM reaches **63.6%** on MLVU and **50.0%** on our 2,400-question **GlanceVU**, showing that substantial video-understanding capability survives rather than collapses at 2 tokens/frame.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.