Promoting Determinism Physics-Conditioned Video Generation via Convergent Flow
Abstract
Physics-conditioned video generation is emerging as a promising paradigm for learning world models that can predict physically plausible futures from initial observations and explicit physical constraints. A distinctive property of these tasks is that the conditioning signals—such as the initial scene, object states, velocities, masses, and interaction parameters—often strongly constrain the future evolution, leading to conditional target distributions that are substantially more concentrated than those in generic video generation. Nevertheless, existing generative approaches typically retain a condition-independent standard Gaussian source. We argue that this default choice is poorly matched to strongly conditioned generation in two respects: its scale can be substantially misaligned with the remaining conditional variation, and its unstructured randomness fails to exploit information already available in the conditioning signal, imposing unnecessary learning burden on the generative flow. To address these limitations, we introduce SCFM, a task-aware source construction for conditional flow matching that adapts the source to both the residual uncertainty and the structural information provided by the condition. By starting the transport from a more informative yet non-degenerate distribution, SCFM enables the flow model to focus its capacity on modeling the remaining conditional dynamics rather than reconstructing structure already specified by the input. Beyond controlled physical supervision, we further introduce VPhysBench, a real-world benchmark for physics-conditioned video prediction, to evaluate whether such source design principles transfer to realistic physical interactions. Experiments on VPhysBench demonstrate that SCFM achieves better performance across multiple distributional metrics with greater learning efficiency, highlighting source distribution design as an important and underexplored dimension of efficient conditional video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.