GroundedHuman: Implicit Interaction Grounding for Dynamic Human–Scene Reconstruction
Abstract
Capturing human motion and surroundings from monocular video requires both temporally coherent reconstruction and plausible human-scene alignment. Existing methods typically trade off the two goals: feed-forward models enable efficient dynamic reconstruction but often yield weaker human-scene consistency, whereas interaction-aware methods often target static settings or incur extra overhead. This work presents GroundedHuman, a feed-forward framework that reconstructs human motion, scene geometry, and camera motion in a shared metric world. Our key insight is to introduce implicit body-scene interaction grounding into feed-forward reconstruction: the body determines where to inspect the scene geometry, producing interaction representations that enable both metric calibration and motion positioning. Specifically, scene geometry is first calibrated with the metric human body via body-aligned scene attention to establish a shared reference for scene depth and camera motion. Then, interactions between body parts and their surrounding scene context are modeled hierarchically for better positioning of human motions. Experiments show that GroundedHuman improves world-space motion estimation and human-scene coherence while reducing floating and penetration artifacts relative to recent state-of-the-art methods on real-world scenarios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.