Analytic Dynamics: Learning Physics-Grounded Representation for Fast Intrinsic Dynamics Inference from Monocular Videos
Abstract
Inferring intrinsic dynamics from observed motion is essential for intelligent agents to reason about and interact with the physical world, yet remains challenging due to the fundamental gap between visual evidence and intrinsic dynamics. Existing methods either rely on costly per-scene optimization, limiting efficiency and scalability, or directly map visual evidence to intrinsic dynamics, making them prone to learning shortcuts that hinder generalization. To address this challenge, we propose Analytic Dynamics, a feed-forward framework that introduces a physics-grounded dynamics representation to bridge monocular videos and intrinsic dynamics. Specifically, we leverage privileged physical states available in simulation, including positions, displacements, and deformation gradients, which explicitly characterize object dynamics, to learn a structured dynamics representation that is difficult to discover from visual observations alone. By aligning video representations with this learned space, we equip video models with a physics-grounded inductive bias, guiding them to capture dynamics-relevant patterns for intrinsic dynamics inference. To support this research, we develop a dynamics data generation pipeline and benchmark comprising paired physical state trajectories, rendered videos, and ground-truth material labels. Extensive experiments demonstrate that Analytic Dynamics outperforms existing baselines, achieves strong generalization, and requires only 3.83 ms per inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.