acceptodds
Under review as a conference paper at ICLR 2027

PhyMem: Retrieved Predictive Memories for Physically Plausible Video Generation

Abstract

Text-to-video diffusion models can synthesize visually compelling clips, yet they remain unreliable when prompts require coherent motion, causal interactions, or physically plausible temporal evolution. A key limitation is that text specifies what should happen but provides little evidence about how the scene should physically unfold. We study whether predictive video representations retrieved from real videos can provide external guidance for physically plausible generation with a frozen Transformer-based latent video diffusion backbone. We introduce PhyMem, a plug-in memory framework that (i) retrieves prompt-relevant reference clips, (ii) instantiates world-model memories with predictive representations extracted by a frozen V-JEPA2 encoder, (iii) compresses the resulting spatiotemporal tokens into compact memory tokens, and (iv) injects them into the self-attention layers of a frozen video diffusion backbone through a lightweight trainable memory module, without finetuning the backbone. Training updates only a lightweight memory interface, with M trainable parameters and approximately K training videos, while the video diffusion backbone, text encoder, and generator VAE remain frozen. Across physics-focused evaluations, PhyMem improves physical plausibility and physical commonsense under matched inference settings while largely preserving semantic adherence. Diagnostic ablations further indicate that the gains depend on representation source, retrieval relevance, and retrieval budget: V-JEPA2 memory outperforms native VAE memory, prompt-relevant references are more effective than shuffled ones, and retrieval performance peaks at an intermediate budget. These results suggest that retrieved predictive memories offer a parameter-efficient route toward improving physical plausibility in video generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.