acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Dense Risk Prediction for Consistent Feed-Forward 3D Reconstruction

Abstract

Recent feed-forward 3D reconstruction models reconstruct dense 3D geometry from unconstrained multi-view images in a single forward pass. As a 3D foundation model, VGGT is increasingly used as a geometric prior in visual SLAM, dynamic scene reconstruction, and novel view synthesis, where VGGT confidence is crucial for point filtering and optimization weighting. However, VGGT confidence is learned jointly with geometric regression rather than directly supervised for depth accuracy, allowing high-confidence regions to contain severe depth errors. Such errors can survive filtering, receive large weights in optimization, and destabilize downstream tasks. We observe that the failures can be captured from VGGT outputs. Depth and point maps can reveal internal self-consistency of VGGT, and multi-view depth maps can capture geometric consistency across views. Motivated by this observation, we propose VERA, a VGGT ERror-Aware dense risk prediction network that adaptively fuses two confidence maps and two consistency signals to predict a dense risk map of depth errors. Trained on nine datasets and evaluated on five sequence-disjoint test sets, VERA substantially improves depth error prediction metrics, reducing AUSE by 30.1% while improving AUROC by 28.6% and AUPR-Error by 81.3% over the best raw signal baseline. When used as an optimization weight, VERA reduces VGGT-Long odometry ATE RMSE by 13.7% and reduces Chamfer distance after confidence-weighted pose refinement by 16.6%, all without retraining or modifying the frozen foundation model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.