acceptodds
Under review as a conference paper at ICLR 2027

VisaRAG: Retrieval-Augmented Dynamics Memory for Future Video Generation

Abstract

Future video generation is ambiguous: visually similar initial states can evolve toward different outcomes, even under the same broad action description. Recent retrieval-augmented methods improve motion realism, physical plausibility, and temporal consistency, but using retrieved experiences to distinguish between such outcomes remains less explored. We propose VisaRAG, a framework that uses observed action-outcome transitions as external dynamics memory for generating video continuations. Given an observed video prefix, VisaRAG retrieves relevant experiences, encodes their temporal changes while suppressing static appearance, and injects the resulting memories through lightweight motion branches into a frozen video generator. A counterfactual learning objective trains the memory-conditioning components to favor memories consistent with the target outcome over alternatives, with the target video and all other generation conditions held fixed. To measure whether retrieved dynamics steer the generated outcome, we further introduce VisaBench, a benchmark that pairs visually similar interactions leading to different outcomes with same-outcome controls, and that conditions every sample on shared, outcome-neutral text so the prompt cannot reveal the target outcome. Across multiple benchmarks, VisaRAG consistently improves both action correctness and the quality of the resulting motion, demonstrating more accurate and realistic future video generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.