Internalizing Rigid-Body Dynamics for Video Generation via Force-Guided Motion Priors
Abstract
Controllable video generation has achieved substantial progress in specifying appearance and motion, yet existing approaches typically rely on dense trajectory, pose, or structural signals that directly prescribe how objects should move. Such controls provide limited understanding of the physical causes underlying motion and can encourage generative models to imitate predefined trajectories rather than reason about dynamics. We introduce ForceGen, a physics-aware video generation framework that conditions motion generation on sparse external-force information. Instead of explicitly specifying future object trajectories, ForceGen represents physical intervention through force-related conditions and learns to translate these signals into plausible spatiotemporal dynamics. To bridge the gap between scalar physical parameters and visual motion, we introduce a Physics-Aware Motion Adapter (PAMA) that integrates visual, semantic, and force information to construct a motion prior for the video generation process. This design enables the model to capture force-induced rigid-body behaviors, including translation, rotation, and their coupled dynamics, while reducing dependence on dense frame-wise motion guidance. Experiments across diverse rigid-body interaction scenarios evaluate both generation quality and physical controllability, with particular emphasis on whether generated motion responds consistently to variations in force conditions. Our results investigate the potential of sparse physical interventions as an alternative control paradigm for physically grounded and controllable video generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.