acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Trust Depth: Scale-Wise Latent Reliability Routing for RGB-D Salient Object Detection

Abstract

Depth benefits RGB-D salient object detection only when its geometric evidence is reliable and relevant to the target; noisy, incomplete, or misaligned depth can instead cause negative cross-modal transfer. We formulate RGB-D fusion as inference of a scale-wise latent reliability variable that controls how the modalities interact at each representation level. Our framework, MPLoRA-MoE, combines shared semantic encoding, modality-specific low-rank adaptation, and four heterogeneous fusion experts. At each scale, an independent router predicts a soft expert distribution from pre-fusion RGB and depth features, allowing depth evidence to be preserved, attenuated, or reconciled without quality or profile annotations. The routed hierarchy is decoded through a single edge-semantic pathway, so conditional computation changes the fused representation rather than selecting among multiple outputs. Across eight RGB-D SOD benchmarks, MPLoRA-MoE achieves state-of-the-art performance on multiple datasets and remains competitive on the others. Ablations further show that the gains depend on reliability-oriented expert specialization, while moderate routing sharpness and compact modality-specific adaptation are sufficient. These results support learning when and how strongly depth should contribute instead of treating it as uniformly trustworthy. Code will be released upon publication.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.