ReconAgent: An Agentic Framework for Evidence-Coupled Sparse-View Indoor Surface Reconstruction
Abstract
Sparse-view indoor surface reconstruction is inherently under-constrained: generative priors can supplement missing observations, but their predictions may also introduce unsupported geometry. Accordingly, we formulate reliable generative reconstruction as two coupled decisions: where to acquire missing evidence and which generated evidence to admit. To this end, we introduce ReconAgent, a closed-loop agentic framework built around a shared scene state, designed to jointly coordinate evidence acquisition and admission. Given the shared scene state, a gain-aware planner selects complementary virtual viewpoints based on viewing-sector coverage and missing-area rate, while a reconstruction-specific vision-language program then converts each selected viewpoint into a structured generation request. The generator synthesizes candidate RGB views with corresponding depth and normal hypotheses, which are subsequently validated by the arbiter against real-view evidence and the current reconstruction before being admitted as valid geometric evidence. Even after admission, generated geometry is integrated only into regions lacking support from real observations, preventing it from overriding observation-supported structure. The admitted evidence updates a canonical gaussian scene, whose refreshed state in turn guides subsequent planning toward unresolved regions. Under the five-view protocol, ReconAgent improves the F-score over the MAtCha baseline by 6.78 points on Replica and 6.54 points on ScanNet++, achieving the best performance on all three geometric metrics while attaining the best or tied-best novel-view synthesis results among the compared methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.