acceptodds
Under review as a conference paper at ICLR 2027

Seeing the Small in the Long: Resolving Where as Well as When for Tiny Evidence in Long Videos

Abstract

Agentic long-video systems have learned to search the timeline, and left one regime unresolved: seeing the small in the long. Every one we evaluate exposes a single tool signature, given a time interval return more frames, so it re-samples when the model looks and never where in the frame. That is insufficient when the queried detail is illegible at the resolution those frames arrive in. To measure that regime we build FoveaBench: 1,044 open-ended questions over 454 videos whose answers turn on a target occupying a median 0.31% of the frame of a median 23-minute video. An item is admitted only if a gate model answers wrongly from a whole-frame span near the evidence moment and correctly from a native-resolution view of it. What the suite exposes is an inversion rather than uniformly low scores: the tool-augmented systems we evaluate score below the tool-free model they are built on, though all of them call their tool on every question. An oracle study over the same items prices what focus is worth, and corrects an intuition: the anchoring interval at native resolution is worth 4.2%; the ground-truth box cropped out is worth less than the frame it was cut from; that same box rendered as a tracked, dilated, context-preserving view is worth 19.6%. That last view changes several things at once, so it prices the rendering pipeline as a bundle rather than any ingredient of it. We then build FoveaSeek, which resolves when first and where in the frame second: it commits to an interval together with an object hypothesis, localises that hypothesis itself as a policy-emitted box rather than by a detector call, so where to look is an action that rollout-level reward can reach, and answers from native-resolution close-ups of the propagated tube. Fine-tuning an 8B backbone reaches 25.2%, against 18.2% for the strongest evaluated baseline; a single-seed reinforcement-learning run under a temporal process reward reaches 31.1%. FoveaSeek demonstrates that the regime is addressable; which ingredient of the rendering bundle carries the gain is left to future measurement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.