Force-Conditioned Video Generation by Decoupling Motion-State Prediction and Video Synthesis
Abstract
Interacting with the physical world using video diffusion models requires modeling how objects move under force control. Existing methods often map force controls directly to video. Without an explicit motion target, they learn force responses only through per-frame visual appearance. Appearance-based prediction may fail to capture the underlying physics, as different motion directions or strengths can appear similar at small motion amplitudes. Consequently, a generated video may look realistic without following the intended control. To improve physical plausibility and controllability, we propose FoMo, a generation scheme that separates motion-state prediction from video synthesis. Specifically, given an initial frame, relative force control, and initial velocity, instead of directly synthesizing the output video, FoMo first predicts initial-frame-indexed velocity and visibility maps. Velocity explicitly represents direction and magnitude, making motion changes direct training targets. In the next step, FoMo integrates the velocities into trajectories and passes them with the visibility maps to a pretrained trajectory-conditioned video generator. This design has two advantages: (1) direct motion supervision and (2) a generator focused only on visual synthesis. We additionally derive relative 2D controls from motion changes in real videos and use semantic physical priors for context. These controls do not represent exact 3D forces. Controlled ablations support explicit velocity prediction and real-video training. In a pairwise study with 100 participants, FoMo is preferred over Goal Force fine-tuned on the same training dataset.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.