PhysGuide: Physics-Principles-Guided Learning for Physically Plausible Video Generation
Abstract
Text-to-video (T2V) models have achieved remarkable visual fidelity, yet their generated videos still frequently violate physical laws. Preference-based post-training provides a promising way to improve physical plausibility, but existing approaches often represent physics supervision using a scalar signal, which collapses the principle-wise structure of physical behavior and provides limited information about how preferred and rejected videos differ physically. We introduce PhysGuide, a physics principles-guided preference optimization algorithm for physically plausible video generation. We construct a set of physics principle vectors from real-world demonstrations and project preferred and rejected videos into this representation space. PhysGuide uses their principle-wise differences to independently modulate the preferred and rejected policy rewards, producing physics-guided rewards along individual physical dimensions. We further introduce a physics-violation divergence penalty that captures the overall magnitude of physical disagreement between the pair, so that more severe violations produce a stronger corrective signal during optimization. We evaluate PhysGuide on PhyGenBench and VideoPhy and observe consistent improvements over the baselines, while preserving general video quality on VBench. Ablations further show that the principle-specific physics-guided rewards and physics-violation divergence penalty contribute complementary improvements to physical plausibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.