acceptodds
Under review as a conference paper at ICLR 2027

HER: Human–Environment Reconstruction via Visual Geometry-Mediated Reasoning

Abstract

Recent monocular HMR methods have made strong progress in world-space motion recovery, yet robust global trajectory estimation remains challenging under monocular ambiguity. Feedforward 3D reconstruction offers complementary spatial cues by recovering camera motion and the geometry of the moving human. We introduce HER, which leverages the moving-human geometry recovered by feedforward 3D reconstruction as a spatial bridge between parametric body prediction and the reconstructed world. HER builds on a backbone-agnostic feedforward geometry frontend to recover spatially consistent geometry and camera trajectories over long videos. For human reconstruction, we freeze a pretrained image HMR encoder and learn a temporal human representation that jointly supports SMPL-X body and explicit structural joint prediction. We further learn a geometry-mediated spatial reasoner that couples this representation with reconstructed human geometry to recover temporally consistent global translation while preserving high-quality pose and shape. Human body structure provides an additional scale reference for establishing a common world coordinate system. A lightweight world-space optimizer further improves temporal and contact consistency. Experiments on EMDB2 and RICH show that HER outperforms state-of-the-art methods in world-space human reconstruction while maintaining strong local accuracy, with robust performance on challenging in-the-wild videos.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.