acceptodds
Under review as a conference paper at ICLR 2027

NBS: No Bias Stereo

Abstract

Stereo reconstruction is one of the last remaining computer vision tasks where state-of-the-art methods still rely on heavy, stereo-specific architectural inductive biases. Although it has been demonstrated that the task can be solved using general-purpose methods, such methods have since fallen behind specialized models, and it is widely believed that inductive biases in stereo are strictly necessary for both high-quality results and computational efficiency. We challenge this paradigm. In this paper, we demonstrate that both state-of-the-art accuracy and superior runtime efficiency are achievable with a model devoid of stereo-specific architectural inductive biases, relying instead on a simple, end-to-end Vision Transformer. By training on large-scale synthetic corpora, we show that a data-driven approach can surpass explicitly engineered geometry: our model ranks first on the ETH3D benchmark, leads on most metrics across four other benchmarks, and runs 4 faster than the previous best method. These results suggest that stereo-specific inductive biases are no longer a prerequisite for stereo matching, paving the way for scalable, general-purpose models for stereo reconstruction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.