DreamSeek: Imagination-Guided Evidence Acquisition for Long-Video Question Answering
Abstract
Finding decisive evidence in a long video can require processing many irrelevant frames. Frame selection reduces the answerer's input, while relevance-guided search uses sparse observations to locate promising regions. However, these observations also provide temporal context for predicting unread visual features, offering guidance beyond measured relevance. We introduce DreamSeek, an imagination-guided approach whose question-agnostic temporal world model predicts unread visual features, enabling relevance assessment before real acquisition. New observations refine the predicted timeline and provide relevance corrections that persist across search rounds. This process enriches search guidance through low-cost latent computation, reducing expensive decoding and encoding. With a shared answerer, DreamSeek achieves the highest mean accuracy on three of five long-video QA benchmarks and up to 5.1 evidence-acquisition speedups over strong frame-selection baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.