Along and Across the Ray: Post-Optimization for Feed-Forward Reconstruction
Abstract
Feed-forward 3D reconstruction models typically process long videos in windows and align their predictions across windows, yet reconstructions of the same surface can still form offset layers and produce surface ghosting. We show that these errors decompose into transverse and radial components in viewing-ray coordinates, where the transverse component is determined by the cameras and ray directions, while with fixed cameras, projection-preserving corrections can only adjust depth along the rays. Based on this structure, we propose RayCon, a training-free global-to-local post-optimization framework. Globally, RayCon exploits the dependence of reconstruction errors on window boundaries through staggered dual-window inference and fuses two trajectories with complementary errors. Locally, with the cameras fixed, Temporal–Range Consistency Filtering removes inconsistent predictions, while Surface-Consensus Refinement infers surface membership from multi-view consistency and refines depths along the viewing rays. Without retraining the backbone, RayCon reduces mean ATE on KITTI from 31.91m to 9.59m, achieves the lowest mean ATE among the compared methods on KITTI and Oxford Spires, and reduces local Chamfer distance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.