Mono4D: Lifting Single-Image 3D Gaussian Reconstruction to Dynamic Decomposed 4D Scenes from Monocular Videos
Abstract
Monocular 4D scene reconstruction is a challenging yet important task for immersive 3D applications. Existing optimization-based methods require multi-view data and costly per-scene optimization, while camera-controlled video generation methods lack explicit 3D representations and demand repeated inference for new camera trajectories; yet recent feedforward 3D reconstruction models perform well on static scenes but suffer from severe geometric and inter-frame inconsistencies when handling dynamic scenes. To address these challenges, we introduce Mono4D, a framework built upon a geometry-aligned static–dynamic decoupled reconstruction paradigm that lifts single-image feedforward 3DGS models to monocular videos of dynamic scenes without additional 4D training or per-scene optimization, enabling efficient nearby novel view rendering. Given a monocular video, Mono4D first estimates camera poses and dynamic masks, then aligns frame-wise 3DGS predictions to a shared world coordinate system. Reliable cross-frame background Gaussians are fused into a globally consistent static representation, while dynamic objects are modeled as temporally indexed Gaussian sequences. To enhance static-dynamic separation, we introduce a boundary-aware mask refinement and Gaussian assignment strategy to mitigate static-dynamic leakage and edge artifacts near fine structures. We further propose a spatiotemporal Gaussian compression scheme leveraging spatial sparsity and temporal redundancy through Gaussian pruning, cluster-level motion encoding, and residual coding for a compact 4D scene representation. Extensive experiments on novel view synthesis benchmarks demonstrate that Mono4D achieves superior performance on monocular 4D scene reconstruction compared with existing approaches, in terms of both visual fidelity and inference efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.