acceptodds
Under review as a conference paper at ICLR 2027

Compressed Sight, Selective Insight: Extreme Long Video Language Models That Reward Their Own Gaze

Abstract

Long-video understanding is constrained by a visual-budget problem: a video may span thousands of frames, but only a small fraction can be given to a language model at high spatial resolution. Existing frame-selection methods retain a few frames and discard the rest, while uniform compression preserves the timeline but allocates the same low resolution to every frame. We introduce Compressed Sight, Selective Insight (CSSI), an efficient video language model that keeps every sampled frame in the input: a small set of selected frames receives a grid of spatial detail tokens, and every other frame is represented by one average-pooled token. Which frames receive detail is decided by a query-conditioned selection policy that reads all frames in one linear-time pass and then selects frames one at a time, each choice conditioned on the frames already chosen, so the budget is spent on complementary evidence rather than repeated views of the same scene. The policy is trained with reinforcement learning from a reward computed by the model itself, the margin by which the model separates the correct answer from the strongest incorrect option when reading the resulting input, so training needs no external reward model and no frame-level labels. Across Video-MME, LongVideoBench, MLVU, EgoSchema, and LVBench, CSSI outperforms recent efficiency-oriented video LLMs, with its largest gains on long videos and on object-reasoning questions. CSSI is also the most memory and latency-efficient method at large frame counts, scaling to 16K-frame videos where competing methods run out of memory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.