GG4D: Geometry-Guided Dynamic Novel-View Synthesis
Abstract
Dynamic novel-view synthesis lets us freely explore a captured moving scene from any viewpoint, enabling free-viewpoint video and immersive VR and AR. To this end, feedforward methods synthesize novel views in a single forward pass without per-scene optimization. Among them, one line predicts explicit 3D representations such as 3D Gaussians, while another directly predicts target views from input images using a learned implicit renderer. Yet explicit methods are limited in fidelity by geometry errors and their fixed rendering process, and implicit renderers without shared 3D geometry struggle to maintain multi-view consistency. In this paper, we present GG4D, a feedforward framework for dynamic novel-view synthesis that combines explicit scene geometry with implicit image synthesis to achieve both multi-view consistency and high rendering fidelity. First, to aggregate geometry across frames without mixing dynamic regions from different timesteps, we propose an efficient segmentation module that separates static and dynamic regions without off-the-shelf optical flow or segmentation models. Using this separation, we construct a 4D point cloud for explicit 3D representation that guides novel-view synthesis. Second, to enable high-fidelity rendering in real time, we propose Large Chunk Attentive Test-Time Training (LaCAT), which effectively encodes input observations into a compact scene memory through attention. Our implicit renderer, built upon it, summarizes the input video once and reuses this information to synthesize high-quality novel views in real time. Extensive experiments demonstrate that \ours achieves state-of-the-art quality and multi-view consistency on both monocular and multi-view dynamic scene benchmarks, taking a step toward immersive applications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.