LEAP: Latent Encoder Alignment for Physics Rewards in Video Generation
Abstract
Reward Feedback Learning is a powerful paradigm for generative modeling, yet its extension to Text-to-Video (T2V) generation is hindered by intractable computational overhead and the scarcity of differentiable reward signals. Although mapping rewards to the latent space reduces computational burdens, existing latent rewards rely heavily on human preference datasets, inherently lacking the rich physical and structural priors of foundational models. In this work, we introduce a novel framework that constructs a latent physics reward by extracting and leveraging priors from a latent world model. We overcome two fundamental hurdles to achieve this integration: 1) developing an alignment mechanism with a lightweight Aligner to connect the generative VAE latent space with the target prior embedding space, and 2) adapting the feedback learning algorithm to stably incorporate these latent priors during training. Our method is the first to demonstrate the viability of projecting T2V latent representations into the V-JEPA 2 vision embedding space for enforcing physical plausibility in generated videos. Empirical results on the VideoPhy and VideoPhy2 datasets validate that our lightweight Aligner helps latent reward substantially improves the physical realism and temporal consistency of generated videos, while training within 12.9GB, where backpropagating the same reward through pixel space would require hundreds of GB.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.