Beyond the Visible: Learning Dense 4D Motion in Contact-Rich Deformable Worlds
Abstract
Predicting 4D dynamics of deformable objects from vision is limited by datasets lacking dense volumetric motion labels and by point-based methods that cannot predict beyond sensor-observed surfaces. We introduce A rena of D eformable dY namics in Temporal Object N exus (ADYTON), a large-scale physics-simulated benchmark with ground-truth material-point trajectories throughout the full 3D volume across ten categories of contact-rich deformable interactions, with systematic generalization tests. We propose D eformable dynamics Evolution Learned via vision-conditioned Predictive H igh-fidelity Implicit fields (DELPHI), a vision-conditioned implicit displacement field that decouples spatial queries from input geometry, enabling amodal prediction from monocular RGB-D. On ADYTON, DELPHI matches the strongest baselines on visible surfaces while extending to unobserved regions with negligible degradation, whereas baselines collapse once denied the privileged correspondence that real sensors cannot provide. These results suggest continuous field representations are a more principled foundation for learned 4D dynamics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.