Pano123: Single-Panorama to Explorable 3D Scene Generation
Abstract
Generating an explorable 3D scene from a single equirectangular panorama remains underconstrained because camera translation reveals surfaces occluded in the input. These disoccluded regions contain no observed geometry or appearance, making their completion the key bottleneck. We present Pano123, a feed-forward framework that predicts a complete 3D Gaussian Splatting scene from a single panorama in one forward pass, without per-scene optimization. Our key insight is that disocclusion should be treated as a sparse, geometrically structured completion problem rather than as dense scene hallucination. Disoccluded regions typically occupy only a small portion of a target view and arise predictably near occlusion boundaries. They can therefore be localized first and completed only where necessary. Accordingly, Pano123 comprises three stages: a sphere-aware Gaussian initializer that accounts for the non-uniform sampling of equirectangular projection; an annotation-free query generator trained with cross-view reprojection pseudo-labels to localize disocclusions and sample candidate points behind visible surfaces; and a context-grounded decoder that synthesizes new Gaussians from neighboring attribute priors and bounded residuals, while preserving visible Gaussians. Experiments on indoor benchmarks show that we outperform prior feed-forward methods, with the largest gains under large camera translations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.