acceptodds
Under review as a conference paper at ICLR 2027

Gaussian for Gaussian: Unifying Feature Representation and Rendering in Feed-Forward Driving Scene Reconstruction

Abstract

Vehicle-mounted surround-view cameras have limited overlap, making it difficult to exploit cross-view information while preserving local detail and reconstructing complete novel views. Existing methods address this challenge through cross-view matching or dense spatial grids, but face limited correspondence support or a resolution–memory trade-off. We propose Gaussian for Gaussian, a feed-forward framework that uses cross-camera features as spatially sparse supplements to local features before constructing a separate renderable scene. Using estimated metric depth, we lift multi-scale image features into Feature Gaussians in a shared 3D space. We splat these Gaussians into the source views and fuse the projected features with local image features, retaining pixel-aligned detail alongside geometrically aligned support. To turn the fused features into high-quality novel views, we further predict independent Render Gaussians with layered geometry and learnable local offsets, providing geometric flexibility beyond the initial depth surfaces. Under known ego-motion, our complete system achieves 28.803 dB PSNR on nuScenes, surpassing the DrivingForward baseline by 7.136 dB, and transfers to PandaSet without fine-tuning with a 4.611 dB advantage. Qualitative and regional analyses show improved reconstruction in newly revealed regions and multi-view overlap regions, while feature aggregation yields larger gains where cross-camera support is stronger.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.