ForceVid3D: Physical Video Generation with Camera-Space 3D Force Control
Abstract
Force-controlled video generation allows users to specify external forces while the model generates scene-dependent responses. Existing direct force-control methods typically rely on image-plane 2D representations, which do not fully express camera-space 3D forces, while 2D observations provide limited cues about responsive regions and their local geometry. We introduce **ForceVid3D**, a framework for generating videos from an initial image and an explicit camera-space 3D force. Our Projective 3D Force Map exposes the image-space geometry of camera-space 3D forces by combining force support, direction, and magnitude with projection cues. We further introduce Depth-Grounded Response Affinity Supervision (D-RAS), which combines response-affinity and local relative-depth objectives to capture responsive regions and their geometric context. To support learning and evaluation across diverse physical responses, we construct ForceVid3D-Dataset, a 60K-pair force–video dataset spanning seven response groups, and introduce ForceVid3D-Bench for force-to-response evaluation. Under comparable 2D force-control settings, ForceVid3D achieves consistent improvements in force adherence and physical consistency over prior force-conditioned methods while maintaining competitive visual quality. Dedicated 3D evaluations further demonstrate camera-space 3D force control, while controlled ablations support the contributions of the projective representation and D-RAS.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.