A Simple Transformer Backend for Feed-Forward 3D Reconstruction
Abstract
Recent feed-forward 3D reconstruction methods such as DUSt3R and VGGT predict pointmaps and camera poses directly from images, yet their internal representations remain image-aligned and lack explicit cross-view geometric reasoning. AMB3R closes this gap by appending a sparse-voxel 3D backend to a frozen frontend, attributing its accuracy gains to the spatial compactness of voxelization. We hypothesize that a much simpler ingredient suffices: aggregating multi-view geometric features and injecting the refined features back into the frontend. We accordingly replace the voxel backend with a deliberately minimal alternative: a single pass of global self-attention over per-pixel geometric tokens, injected back into the frozen decoder through zero-initialized adapters, with no voxel grid, no space-filling-curve serialization, and no -nearest-neighbor interpolation. Despite its simplicity, the proposed backend is comparable to a voxel backend trained under matched conditions across zero-shot monocular depth, camera pose, multi-view depth, multi-view 3D reconstruction, and video depth. Controlled ablations on 3D positional encoding, the backend-to-frontend injection operator, and iterative refinement suggest that a single feed-forward pass with minimal design machinery is sufficient. Finally, on three additional frontends (, Déjà View, and Depth Anything 3), our backend performs comparably to or better than a voxel backend trained on the same frontend, suggesting that the mechanism studied in this paper transfers beyond VGGT.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.