DiPhy: Distilling Transferable Physical Knowledge for Physics-Aware Video Generation
Abstract
Recent generative video models have achieved impressive visual quality and increasingly realistic video synthesis, yet they still struggle to generate videos that obey plausible physical dynamics. Existing approaches typically incorporate physical information through textual descriptions, trajectory constraints, or additional conditioning signals, whereas how physical knowledge can be effectively mined and transferred to video generation has not yet been sufficiently investigated. Facing this challenge, we propose DiPhy, which distills transferable physical knowledge into a diffusion model for physically plausible video generation. Specifically, a JEPA-based teacher first mines physics-related cues from videos under semantic guidance. These cues are organized into a structured physical codebook that serves as a compact shared space for transferable physical knowledge, which is then distilled into a student module for text-only physical reasoning without video access. The inferred physical codes are subsequently injected into the diffusion backbone through a lightweight physical branch, providing explicit physical guidance during generation. Extensive experiments on physics-oriented video benchmarks demonstrate that our DiPhy improves physical consistency while preserving semantic alignment and overall video quality, with physical commonsense scores improving by approximately 9% and Joint performance by over 10% on VideoPhy2. Moreover, our DiPhy enables efficient video generation with relatively short inference time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.