acceptodds
Under review as a conference paper at ICLR 2027

Scene and Human in One World: Reconstruction in a Feedforward Pass

Abstract

Reconstructing humans in dynamic scenes from monocular video remains challenging due to scale ambiguity and human–scene misalignment. We observe that a monocular point map is defined only up to a single global scale, which scene geometry alone cannot resolve, whereas the parametric human body has a statistically well-constrained real-world size. We introduce SHOW, a mask-promptable feed-forward framework that reconstructs a designated person together with the surrounding scene in one shared metric space, in which the human body acts as the metric reference that determines the scale of the normalized point map. Mask prompting and auxiliary DensePose supervision make the geometry features discriminative on human regions, and a unified decoder regresses body pose, shape, and global translation together with the scale factor of the scene, so that scale, placement, and scene geometry are optimized as a single system instead of being aligned post hoc. SHOW therefore produces a human reconstruction that is already metrically consistent with the reconstructed scene. On 3DPW, EMDB, and RICH, it improves metric-scale consistency and human–scene alignment over prior joint human–scene feed-forward methods while remaining competitive on global human motion estimation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.