VideoArtGS: Building Digital Twins of Articulated Objects from Monocular Video
Abstract
Building digital twins of articulated objects from monocular video requires jointly recovering part geometry and articulation from entangled camera and object motion. Although 3D point tracks provide motion cues, their noise and lack of explicit part structure complicate articulation learning. To address this problem, we introduce VideoArtGS, a motion-prior-guided framework that turns these tracks into structured initialization for joint geometry-and-motion reconstruction. Our pipeline analyzes trajectories under articulation constraints, filters unreliable tracks, and groups them to initialize part centers, joint parameters, and an articulation-based deformation field. We also design a hybrid center-grid part-assignment module that represents movable parts with learnable centers and the static base with a flexible spatial grid. We then jointly optimize 3D Gaussians and the deformation field using rendering and tracking supervision. VideoArtGS achieves state-of-the-art performance in articulation estimation and mesh reconstruction on iTACO-S, iTACO-L, and VideoArtGS-20 datasets. On VideoArtGS-20, which contains objects with up to nine movable parts, mean joint-axis error decreases from to , a reduction of over two orders of magnitude. Visualizations are available at: https://videoartgs-2026.github.io.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.