acceptodds
Under review as a conference paper at ICLR 2027

Memory-Grounded Video Latent Reasoning

Abstract

Long-video understanding and complex video reasoning require vision–language models (VLMs) to identify sparse visual cues amid redundant content and integrate evidence across distant moments through complex reasoning. Textual chain-of-thought (CoT) and agent-based approaches incur high inference costs; extended textual reasoning can lose focus on visual evidence, while repeated local inspection can reinforce errors. Although image-based visual latent reasoning reduces reliance on lengthy textual descriptions, we observed that directly transferring it to videos degrades performance, revealing visual forgetting and representation shift. We propose Memory-Grounded Video Latent Reasoning (MVLR), which interleaves textual reasoning with continuous visual latent reasoning grounded in adaptive visual memory. Guided by textual CoT, the memory preserves keyframe detail, compresses complementary context, and supplies relevant evidence at every latent step. Two-stage supervised training combines CoT and answer supervision with retrieval diversity regularization, learning latent representations without explicit visual-feature targets. Counterfactual-constrained reinforcement learning further assigns contrastive credit by comparing the model's support for correct answers under factual memory with a frozen reference model's support for the same answers under counterfactual memory. MVLR achieves state-of-the-art performance across 5 benchmarks spanning long-video understanding and complex video reasoning, while delivering 2.63x the inference throughput of strong baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.