Invalid Depth Is Not Missing Depth
Abstract
A missing depth measurement is not a measurement of zero. Yet common RGB-D unseen object instance segmentation pipelines encode invalid pixels as XYZ = (0, 0, 0), placing them at the optical centre. On OCID and OSD, 11–24% of pixels lack depth, whereas the synthetic training data are complete. Controlled comparisons and within-model ablations identify invalidity encoding as a substantial confound in robustness evaluation: UCN without validity filtering drops to zero Object-F after just 0.5% additional i.i.d. pixel removal, while the tested image-depth baseline remains stable. We further show that i.i.d. removal poorly matches the spatial structure of observed sensor failures, and introduce a corruption generator calibrated to each sensor's missing-depth statistics. Our repair, ProxFill, copies the XYZ triplet of the nearest valid pixel into an invalid pixel only when their image-space distance is at most τ. With τ fixed on a disjoint development split, this training-free rule improves the official UCN+ checkpoint by 2.8 Object-F and 6.9 Boundary-F points on OSD and transfers unchanged to MSMFormer. Under calibrated degradation, corrected UCN and MSMFormer remain within one Object-F point of their clean-input scores, while the published pipelines lose 14–16 points. Replacing the depth-validity filter with ProxFill also reduces measured UCN processing latency by 25%, without retraining or additional learnable parameters. These findings show that a substantial part of the measured robustness gap between RGB-D methods can arise from how invalid depth is represented.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.