ForceStereo: Observation Assimilation through State Evolution for Stereo Matching
Abstract
Modern stereo matching methods achieve strong overall accuracy by combining stereo matching cues, contextual features, and monocular depth priors. However, they rarely assess how reliable different image regions are in open world scenes, so overall metrics and visual results can obscure large errors in regions with unreliable observations. These hidden errors may later surface as noise or artifacts in downstream 3D pipelines. We propose ForceStereo, a depth state inference framework based on observation calibration that reformulates disparity estimation as state evolution guided by reliability. ForceStereo first uses scene-level context to calibrate ambiguous local disparity hypotheses before state formation, so the model does not commit too early to unreliable matches. It then interprets residuals, hypothesis ambiguity, and structural information as reliability conditions, and uses a Neural ODE to regulate their influence throughout disparity evolution. This inference process controls how unreliable observations affect the evolving disparity state instead of directly turning them into disparity updates, improving robustness and generalization in open world scenes. Extensive experiments establish state-of-the-art performance, including up to 16.0% error reduction on the ETH3D leaderboards and up to 25.3% zero-shot error reduction over prior best results on challenging datasets. ForceStereo ranks 1st on multiple key metrics in both the KITTI 2012 and KITTI 2015 stereo benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.