MoVi: Motion-Centric 4D Dog Reconstruction from Monocular Videos
Abstract
We present **MoVi**, a motion-centric framework for reconstructing 4D dog motion from unconstrained monocular RGB videos for quadruped imitation learning. Existing animal motion acquisition pipelines typically depend on controlled capture setups, specialized sensing, or calibrated recording environments, which limits the scale and diversity of motion data available for downstream learning. MoVi instead targets readily available in-the-wild videos and converts them into retargeting-friendly motion representations. Our framework first estimates per-frame 3D skeletons with a flow-matching-based pose model, then refines them using reprojection constraints and temporal modeling to improve consistency under monocular ambiguity and occlusion. We further introduce a trajectory predictor that recovers global root motion from articulated motion and temporal cues, producing spatially coherent 4D motion rather than appearance-oriented reconstruction. By focusing on articulated pose and global trajectory, MoVi provides smoother and more task-relevant motion representations for quadruped retargeting and imitation learning, offering a scalable way to leverage large collections of ordinary monocular videos as training data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.