Rectify Pose and Capture Ambiguity: 3D Spatial-Aware GS in the Wild
Abstract
Recent 3DGS and NeRF variants have significantly advanced for real-world scenarios. However, they still struggle to capture ambiguous distractors that are visually or semantically similar to static regions and remain susceptible to inaccurate camera poses. We found that these two challenges are inherently coupled, making them difficult to address independently and leading to artifact-prone 3D reconstruction. Motivated by these findings, we introduce PAGauss, a novel framework that jointly addresses these challenges. Specifically, unlike previous works that focus on 2D representations, PAGauss leverages 3D spatial relationships across viewpoints. PAGauss consists of two novel components: a spatial-aware adaptive mask and a spatial-aware camera pose alignment. The spatial-aware adaptive mask leverages 3D multi-view consistency to robustly capture ambiguous distractors regardless of their colors or semantic similarity. The spatial-aware camera pose alignment rectifies imprecise camera poses using a source refinement strategy with distractor-free images. To stabilize camera pose alignment, we introduce a multi-stage training strategy, improving 3D Gaussian alignment. We also create and release the WildArena dataset that includes challenging scenarios with corresponding mask annotations. To our knowledge, WildArena is the first dataset to provide such mask annotations, enabling robustness analysis over baselines. Through extensive experiments, we demonstrate that PAGauss achieves superior results and robustness across in-the-wild datasets. Our source code and dataset will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.