MetaV3RSE: 4D Human-Object-Scene Reconstruction from Monocular Video
Abstract
We introduce , a unified feed-forward framework that reconstructs the human, the manipulated object, and the surrounding scene together in one online pass from monocular video. Existing feed-forward methods address either human-object or human-scene reconstruction, not the coupled setting, so object recovery cannot use the human's scene grounding and the scene-grounded human cannot use the constraints supplied by the object. Methods that do reconstruct all three rely on per-sequence optimization. The difficulty is that the manipulated object is the least visible part of the interaction. The reconstructed human constrains where the object lies but not its precise rigid pose, while visible object evidence is precise but incomplete. We therefore extend the recurrent human-scene state with object and human-object contact representations that produce an interaction-conditioned pose hypothesis and feed interaction back to the human, then correct that hypothesis using visual evidence gathered at canonical points of the known object template. We show that reduces combined human-object reconstruction error by 24% over the strongest feed-forward baseline on BEHAVE, and joint object reasoning also improves the human, reducing WA-MPJPE by over 23% on the unseen EMDB-2 benchmark.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.