LEG4D: Language Embedded Generative 4D Gaussian Splatting
Abstract
Reconstructing complete and semantically meaningful 4D scenes from monocular videos is inherently challenging due to limited observations of moving objects. Existing dynamic reconstruction methods mainly recover visible geometry and appearance, while generative approaches often require costly score-distillation optimization and lack dense semantic representations. We present LEG4D, a Language-Embedded Generative 4D Gaussian Splatting framework for complete, semantic, and editable dynamic scene reconstruction. Given a monocular video and object masks, LEG4D employs a pretrained 3D object generation model to complete partially observed objects as object-level Gaussians. The generated Gaussians are integrated into the observed scene through pose- and geometry-aware alignment, multi-view appearance adaptation, and object replacement. A motion-aware dynamic Gaussian module subsequently optimizes the fused scene while maintaining Gaussian correspondences and trajectories across time. We further distill language-aligned features from observed frames into the dynamic Gaussians, producing a complete, temporally consistent, and language-embedded 4D representation. Experiments demonstrate that LEG4D improves both dynamic scene completeness and language-grounding accuracy over prior approaches. The resulting representation additionally supports object-level querying, localization, removal, and localized editing in dynamic scenes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.