acceptodds
Under review as a conference paper at ICLR 2027

TimeQuest: Scaling Inference-Time Compute in Video-LLMs via Decoupled Temporal Navigation

Abstract

Scaling inference-time compute has improved the performance of Large Language Models (LLMs), yet its application to video understanding is limited by a fundamental perception bottleneck. Existing Video-LLMs primarily scale reasoning in the text space; however, if a model misses brief but critical visual evidence during its initial pass, no amount of subsequent text-based "thinking" can recover it. In this paper, we introduce TimeQuest, a framework that decouples temporal evidence acquisition from reasoning via a novel Navigator–Oracle architecture. The Navigator, trained using Group Relative Policy Optimization (GRPO) with a smooth distance-based temporal reward, iteratively identifies informative video segments. Rather than naively stacking raw high-dimensional visual tokens (which rapidly exhausts context window), the Oracle translates these localized segments into dense textual descriptions. This accumulated context is appended to the initial query, enabling the Oracle to process the original video alongside an evidence-augmented prompt to deduce the final answer. This design fundamentally shifts inference-time compute from raw visual processing to active evidence distillation. Extensive evaluations on standard and temporally grounded Video QA benchmarks demonstrate consistent improvements of up to **+6.0** accuracy points, confirming that active evidence acquisition is a key factor in effective inference-time scaling for video understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.