In-Context Search with Masked Resampling for Long-Video Temporal Grounding
Abstract
Video large language models (Video LLMs) localize moments in short videos with high accuracy, yet struggle when the target lasts a few seconds within hours of video. Existing multi-stage methods for long videos, with or without Video LLMs, search from coarse to fine, but refine within a single contiguous window at a time; if that window misses the target, the refinement is wasted. We instead bridge the coarse and fine stages with an in-context search that carries several candidates from one stage to the next. Since a Video LLM outputs only the window it is most confident about, we propose masked resampling, which hides the frames of each predicted window from the model and lets it answer again, so that a single pass over the video yields several distinct candidates. The candidates are joined into a montage, a sequence of short clips shown with their original timestamps, on which we find that Video LLMs ground without any adaptation. Montages also remove the need to feed a uniformly downsampled whole video to the Video LLM as its coarse input: an embedder instead filters out clearly irrelevant content, and the Video LLM searches a montage of the most relevant clips, which retains the target in most cases at a higher per-frame resolution. The resulting method, MaRS, requires no training and improves three Video LLMs on four long-video benchmarks at the same visual budget, including models already trained on long videos. When long-video training data are available, fine-tuning on montages reduces long-video training to training on short inputs and brings further gains: with MaRS, a fine-tuned 4B model outperforms larger fine-tuned Video LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.