acceptodds
Under review as a conference paper at ICLR 2027

SplatRefiner: Faithful and Efficient Post-Hoc Diffusion Refinement for Sparse-View Novel View Synthesis

Abstract

Sparse-view novel view synthesis requires inferring content that a few input images leave unresolved. Many diffusion-assisted pipelines supply this content through reconstruction–generation cycles along predefined camera trajectories, making novel views costly to obtain. Refining renderings of a fixed Gaussian reconstruction directly into final outputs avoids these cycles. However, these renderings mix useful scene content with artifacts, and without 3D re-optimization, the refiner alone must remain faithful to the observed scene. To address this, we exploit information from both the reference views and the target view, during training and at inference. During training, VGGT features of the reference views condition denoising, and those of the ground-truth target view supervise intermediate representations. At inference, dual-condition guidance strengthens the contributions of the reference views and of the rendering of the target view. Motivated by these, we present SplatRefiner, a post-hoc diffusion refiner that restores renderings at arbitrary requested poses without predefined trajectories. On identical renderings, SplatRefiner outperforms strong baselines in most experiments while in 3 input view settings it runs 6.9 faster end-to-end with comparable PSNR using 2DGS, and 3.7 faster with higher PSNR using FSGS. Trained only on 2DGS renderings from DL3DV-10K, it improves PSNR, LPIPS, and FID over the renderings of 2DGS, 3DGS, and FSGS across three datasets, two unseen during training, with 3, 6, and 9 input views.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.