BRIDGE: Building Reward Optimized Depth-to-Image Data Generation Engine for Monocular Depth Estimation
Abstract
Monocular Depth Estimation (MDE) recovers scene geometry from a single RGB image, yet what a model can learn is capped less by its architecture than by the depth supervision available: sensor-captured RGB-D is metrically accurate but sparse, rendered synthetic RGB-D is dense but domain-shifted, and web-scale pseudo-labels cover broad appearance but inherit teacher boundary noise. We propose BRIDGE, a reward-optimized depth-to-image (D2I) engine that synthesizes over 20M RGB images from diverse source depth maps, each paired with the source depth map used to condition it. We then train a depth model on this data with similarity-guided hybrid supervision: source depth supervises regions where the generated image preserved geometry, and teacher pseudo-labels supervise where it drifted. BRIDGE achieves the strongest zero-shot relative-depth results among compared methods on NYUv2, ScanNet, ETH3D, and Sintel, using 20M generated images against Depth Anything V2's 62M.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.