Velo3R: Motion-Guided Unified Online 4D Human-Scene Reconstruction
Abstract
Online 4D human–scene reconstruction aims to recover human motion, camera trajectory, and scene geometry from a monocular video as it arrives, using only the current and preceding frames and expressing every estimate in a shared world coordinate system. Maintaining coherent motion is especially challenging online, as occlusions and ambiguities between body and camera movement must be handled without access to future frames. These ambiguities can lead to inconsistent frame-to-frame estimates, producing jitter and unreliable global trajectories. We seek an explicit link between successive body states to encourage temporal consistency while allowing new observations to correct the reconstruction. To this end, we introduce Velo3R, which uses explicit motion estimates to refine the latent human representations used for current-frame mesh prediction. Conditioned on temporal tokens and previously reconstructed body states, our conditional velocity field links joint positions, velocities, and accelerations through integration and differentiation. Inspired by physics-informed neural networks, we jointly supervise these quantities and feed field-derived motion features back into human representations before mesh regression. Experiments on EMDB-2 and RICH demonstrate improvements in global motion accuracy and temporal consistency. Compared with Human3R, Velo3R reduces W-MPJPE from 267.9 to 259.6 mm on EMDB-2 and from 184.9 to 170.6mm on RICH. It also reduces world-coordinate jitter by 9.9% and 14.5%, respectively, with lower foot sliding on both datasets. These results highlight the value of explicit motion feedback for temporally coherent online 4D human–scene reconstruction. Our implementation will be made available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.