UniDisp: Learning Unified Hierarchical Stereo Representations for Progressive Disparity Estimation
Abstract
Deep stereo disparity estimation has long followed staged pipelines of feature extraction, correspondence modeling, and disparity refinement. Although effective, such pipelines distribute stereo representation learning and disparity inference across multiple modules, while global ambiguous correspondences can remain difficult to resolve from local matching evidence alone. We propose UniDisp, a unified backbone that estimates disparity from stereo representations. This is motivated by a simple hypothesis: disparity can be treated as an intrinsic geometric attribute encoded in stereo representations, rather than a quantity derived from matching within a modular stereo pipeline. We instantiate this idea with a hierarchical stereo encoder that jointly models cross-view displacement cues across spatial scales, a prediction head that maps multi-level stereo representations directly to disparity, organized into a two-stage framework with structurally identical yet functionally complementary encoder–head stages. The first stage performs initial estimation from the raw stereo input, while the second uses implicit warping to construct disparity-aware inputs for residual refinement. This design unifies stereo representation learning, disparity prediction, and refinement into a compact backbone, enabling disparity to be inferred from the learned stereo representations themselves. Extensive experiments demonstrate consistent cross-domain transfer, reaching state-of-the-art performance on several real-world benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.