Discarding Sensor Depth: Reliability-Aware Geometry Learning for Salient Object Detection
Abstract
RGB-D salient object detection (SOD) complements RGB appearance with scene geometry, but practical depth measurements are often noisy, incomplete, or structurally misleading, making tightly coupled RGB-depth models vulnerable to negative transfer. We propose \method, a depth-free RGB-D SOD framework that learns task-oriented geometry from RGB instead of consuming sensor depth. During training, a frozen foundation depth estimation model transfers dense relative structure, multi-scale feature organization, and boundary cues to an RGB-driven geometry branch, addressing how useful geometry is acquired without dataset depth. Cross-modal enhancement then exchanges appearance and geometry across scales, while reliability-aware fusion suppresses geometric responses that disagree with RGB evidence or are irrelevant to saliency, addressing how the learned geometry should be used. The external teacher is removed after training and sensor depth is never required by the final detector. Across nine RGB-D benchmarks, is best or tied-best in 26 of 36 metric–dataset comparisons, with the clearest gains on cross-dataset benchmarks. Results on four RGB SOD benchmarks further confirm that the learned geometry transfers beyond RGB-D sensor domains. Overall, task-adapted geometry retains the structural benefit of depth while avoiding direct dependence on unreliable measurements.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.