Joint Modeling of Visual and Physical Fields for Consistent Video Generation
Abstract
Text-to-video generation models have achieved remarkable progress in visual fidelity and temporal coherence, yet generated videos can still exhibit implausible motion dynamics and inconsistent interactions. Existing efforts incorporate physical information (e.g., structured motion or geometric knowledge) mainly via external guidance or constraints, failing to directly consider it within the visual generative representation. In this paper, we proposed to employ physical field signals (e.g., velocity, acceleration, and depth) as a quantitative outcome of motion which naturally serves as unified structured representation across different motion scenarios, and then simultanesouly optimize phyical and visual signals in representation space. However, physical field signals is sparse and noisy in real videos, posing siginificant difficulties in the joint optimization. To tackle these issues, we propose a novel VideoPF method that jointly models video data and physical field signals within a unified generative framework, enabling the synthesis of videos with improved motion consistency. Specifically, we first analyze the degree to which generative models capture physical dynamics. Based on this analysis, we design an adaptive physical-aware data pipeline. To facilitate effective fusion of visual and physical information, we introduce a Gated Physical DiT that selectively integrates physical field signals into the visual generation process. To enhance learning stability and encourage physically plausible dynamics, we propose a Physical Consistency Refinement mechanism and dynamically assess the model’s understanding of physical information. Extensive experiments demonstrate the effectiveness of our proposed VideoPF framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.