I Feel How You See: A Dynamics World Model for Predicting Forces and Kinematics for High-touch Robotic Manipulation
Abstract
In contact-rich manipulation, forces and torques are the fundamental drivers of contact dynamics, dictating how objects move and react to manipulation. Yet, most world models today forgo the grounding that comes from predicting plausible force/torque (F/T) traces alongside their predicted robot trajectories. Addressing this gap, we present a Bin Dynamics World Model, BDWM, for robotic bin stowing via space-creating sweep motions. Conditioned solely on pre-contact information (an RGB-D image, item segmentation, and the planned blade sweep), a single flow-matching velocity field jointly generates two latent streams: a physics latent that decodes to the full F/T and end-effector trajectory, and a vision latent that predicts the post-contact scene. Evaluated on real-robot cycles, BDWM reduces trajectory CRPS by 75% (0.009 vs. 0.036) over the strongest generative action-model baseline and trajectory RMSE by 26% (0.025 vs. 0.034 m) over the strongest baseline overall. Lightweight heads on the predicted latents further drive downstream tasks such as the prediction of space created after a sweep, reaching sub-centimeter accuracy compared to a task-specific regression baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.