Beyond the Visible Surface: Geometry-Guided Joint Optimization of Multi-Layer Gaussians for Single-View Scene Reconstruction
Abstract
Feed-forward Gaussian reconstruction predicts a renderable 3D scene from a single image, yet reconstructing content beyond visible surfaces remains challenging. Limited hypotheses along each source ray restrict the representation of newly exposed regions, while depth-based initialization alone provides insufficient geometric guidance for feature learning. We propose GeoJMG, a geometry-guided framework for jointly learning multi-layer Gaussian representations for single-view scene reconstruction. Ray-Stratified Gaussian Modeling (RSGM) constructs multiple depth-ordered surface hypotheses along each source ray and predicts their Gaussian parameters through independent decoder branches over shared multi-scale features. All layers are trained jointly under a common reconstruction objective without predefined visibility roles or layer-specific supervision. To guide this learning process, we derive a compact geometric descriptor from depth, camera-space coordinates, depth gradients, surface normals, and reliability. A lightweight Geometry-Conditioned Residual Adapter (GCRA) integrates these cues into multiple encoder scales through zero-initialized projections, preserving the base features at initialization. The resulting model reconstructs a layered Gaussian scene in a single forward pass without per-scene optimization. Experiments on RealEstate10K, NYUv2, and KITTI demonstrate improved novel-view synthesis quality and effective generalization across indoor and outdoor scenes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.