HiFiNVS: High-Fidelity Novel View Synthesis from Unposed Images
Abstract
Feed-Forward Novel View Synthesis from unposed images has recently benefited from pretrained visual geometry models, which provide strong camera and dense geometry priors for Gaussian reconstruction. Nevertheless, the resulting Gaussian predictions can remain inconsistent across views, while errors in the estimated cameras often lead to ghosting and blurred details. We introduce HiFiNVS, a backbone-agnostic NVS framework that exploits predicted geometry not only for reconstruction, but also to establish explicit cross-view constraints for feature adaptation and camera refinement. Specifically, our Reprojection-Grounded Graph Attention constructs a patch graph by reprojecting 3D points across images and propagates information along geometrically grounded correspondence edges, thereby improving the cross-view consistency of Gaussian predictions. We further introduce a feed-forward camera refiner that updates the predicted cameras using rendering feedback from the initial Gaussian scene. Rerunning our reconstruction once with the refined cameras effectively suppresses rendering artifacts. HiFiNVS keeps the visual geometry backbone frozen and can be instantiated with DA3, π³, and VGGT. Across all three backbones, camera refinement improves both camera accuracy and rendering quality. Experiments on unposed-view benchmarks show that HiFiNVS achieves state-of-the-art NVS performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.