Lightweight Test-Time Geometry Unification for Multi-View Reconstruction
Abstract
Feed‑forward 3D reconstruction models can predict several offset sheets for one physical surface. Reconciling these predictions requires both cross‑view agreement and preservation of local surface structure: independent voxel shifts can narrow layers while introducing cracks between neighboring pixels. We introduce a self‑supervised test‑time refiner that addresses these coupled requirements without target‑scene depth or 3D reference supervision. Closed‑form per‑view similarity alignment first corrects rigid‑plus‑scale drift. A 0.9 M‑parameter residual network then corrects the remaining non‑rigid component of multi‑view geometric inconsistency, using multi‑view image features and geometric self‑supervision with displacement‑field regularization to preserve local structure. Wall‑corner and curved‑edge comparisons show that this second stage brings residual sheets closer together while retaining surface continuity. It reduces cross‑view centroid gaps in all 53 evaluated NRGBD/DTU scene–model cases. On 6 658 patches from six NRGBD scenes, its adjacent‑pixel GT edge‑vector error is lower than full‑strength direct () and smoothed voxel shifts in 94.7% and 87.5% of cases, respectively. The same refiner applies to Pi3 and VGGT‑ without per‑base retraining, improving Accuracy and Normal Consistency across the evaluated object‑level, indoor, and outdoor benchmarks. The complete pipeline takes about 3.7 minutes for a representative 15‑view scene on one A800 GPU. A fixed‑count analysis shows that redundant layers can lower Completeness error while raising Accuracy error, motivating evaluation of surface structure alongside point distances.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.