acceptodds
Under review as a conference paper at ICLR 2027

DiPhy: Distilling Transferable Physical Knowledge for Physics-Aware Video Generation

Abstract

Recent generative video models have achieved impressive visual quality and increasingly realistic video synthesis, yet they still struggle to generate videos that obey plausible physical dynamics. Existing approaches typically incorporate physical information through textual descriptions, trajectory constraints, or additional conditioning signals, whereas how physical knowledge can be effectively mined and transferred to video generation has not yet been sufficiently investigated. Facing this challenge, we propose DiPhy, which distills transferable physical knowledge into a diffusion model for physically plausible video generation. Specifically, a JEPA-based teacher first mines physics-related cues from videos under semantic guidance. These cues are organized into a structured physical codebook that serves as a compact shared space for transferable physical knowledge, which is then distilled into a student module for text-only physical reasoning without video access. The inferred physical codes are subsequently injected into the diffusion backbone through a lightweight physical branch, providing explicit physical guidance during generation. Extensive experiments on physics-oriented video benchmarks demonstrate that our DiPhy improves physical consistency while preserving semantic alignment and overall video quality, with physical commonsense scores improving by approximately 9% and Joint performance by over 10% on VideoPhy2. Moreover, our DiPhy enables efficient video generation with relatively short inference time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.