MoVE-4D: Feed-forward Video-to-4D Generation via 4D Structured Gaussians
Abstract
Generating dynamic 4D assets from monocular videos in a feed-forward manner has rapidly progressed by extending pretrained 3D foundation models to time. However, existing paradigms still treat 3D representations as the atomic unit and introduce temporal modeling externally, either through canonical-shape deformation or per-frame generation coupled by temporal attention. Both designs leave 3D structure unshared at the representation level, limiting motion expressivity in deformation-based methods and leaving a persistent quality–efficiency trade-off in per-frame methods. We propose MoVE-4D, a feed-forward video-to-4D framework built on a 4D-native representation that shares 3D structure across time at the representation level itself. Specifically, we introduce structured 4D Gaussians on motion-occupied voxels, where sparse voxel anchors are reused across frames within a temporal window and time-dependent Gaussian attributes capture local dynamics. MoVE-4D generates this representation in two stages: it first predicts motion-occupied regions through voxel-wise occupancy merging, then populates them with structured 4D Gaussians decoded from a Compact Temporal Structured LATent that aggregates multi-frame visual features anchored at the same voxel. Extensive experiments show that MoVE-4D delivers a Pareto-optimal quality–efficiency trade-off, surpassing existing feed-forward methods in visual quality and spatiotemporal consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.