acceptodds
Under review as a conference paper at ICLR 2027

Look Again: Attention-Guided Temporal Resampling for VideoLLMs

Abstract

Uniform frame sampling can miss the brief events needed to answer questions about long videos. We investigate whether a VideoLLM’s internal attention can guide a second look at relevant moments, even when the initial view does not support a correct answer. We introduce Temporal Attention Resampling (TAR), a training-free method that contrasts question-conditioned attention with a generic reference, retrieves a contiguous temporal interval, and combines densely resampled local frames with global context. On our 502-question MLVU evaluation set, TAR improves Qwen3-VL-8B from 44.42% to 50.20%, outperforming matched 48-frame uniform sampling by 4.38 percentage points and random-window re-sampling by 4.65 points. Accuracy improves across eight benchmarks and two Qwen-family model scales. On the annotated LVBench subset, questions whose retrieved windows overlap evidence gain 6.43 points, compared with 1.19 points for those without overlap. This association is consistent with improved temporal evidence acquisition. TAR requires no parameter updates, temporal supervision, or external retriever.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.