Look Again: Attention-Guided Temporal Resampling for VideoLLMs
Abstract
Uniform frame sampling can miss the brief events needed to answer questions about long videos. We investigate whether a VideoLLM’s internal attention can guide a second look at relevant moments, even when the initial view does not support a correct answer. We introduce Temporal Attention Resampling (TAR), a training-free method that contrasts question-conditioned attention with a generic reference, retrieves a contiguous temporal interval, and combines densely resampled local frames with global context. On our 502-question MLVU evaluation set, TAR improves Qwen3-VL-8B from 44.42% to 50.20%, outperforming matched 48-frame uniform sampling by 4.38 percentage points and random-window re-sampling by 4.65 points. Accuracy improves across eight benchmarks and two Qwen-family model scales. On the annotated LVBench subset, questions whose retrieved windows overlap evidence gain 6.43 points, compared with 1.19 points for those without overlap. This association is consistent with improved temporal evidence acquisition. TAR requires no parameter updates, temporal supervision, or external retriever.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.