S2D-NeRF: Screen and Space Decoupling for Distractor-Free Novel View Synthesis
Abstract
Neural Radiance Fields (NeRFs) achieve photorealistic novel view synthesis, but assume that the scene is consistently observed across all views. In real-world captures, transient objects and the shadows they cast violate this assumption and lead to floaters and ghosting artifacts. Existing methods often identify distractors using semantic priors or heuristic rules, making the separation sensitive to segmentation quality or hand-designed assignment criteria. Instead, we find that cross-view consistency itself provides a strong cue for separating distractors, and our representation is designed to turn this cue into an effective inductive bias. We therefore propose SD-NeRF, which models the static scene with a shared 3D field and distractors with per-image 2D fields directly defined in pixel coordinates. This asymmetric representation drives their separation during joint optimization from the first iteration. Experiments on the NeRF On-the-go, RobustNeRF, and Photo Tourism datasets show competitive or superior rendering quality without semantic priors or heuristic rules. Our code will be publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.