Art2Sim: Learning Physically Executable Full-Body Humanoid Interaction with Articulated Objects from Video Priors
Abstract
Video priors provide a scalable source of whole-body human-object interactions, yet visually coherent trajectories do not guarantee successful physical execution. This gap is especially acute for articulated objects. To open a door or pull a drawer, a humanoid must sustain contact with the correct moving link and coordinate whole-body motion to drive the object's passive joint. To bridge this gap, we introduce Art2Sim, a video-to-physics framework that transforms a monocular interaction video paired with an articulated asset into physically feasible full-body manipulation. Our framework consists of two stages: (1) Contact-Aware Articulated Reconstruction builds a contact-consistent kinematic reference that aligns the hand with the moving link; (2) Articulation-Guided Physics Refinement learns a residual policy on top of a pretrained whole-body tracker, using contact, joint-progress, and mechanical-work rewards to convert the reconstructed 4D HOI reference into feasible interactions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.