WEAVE: Warp Every Angle via 3D Geometry for Reference-Guided Image Completion
Abstract
Reference-guided image completion uses auxiliary views of a scene to reconstruct missing image content faithfully. Existing approaches often require per-scene optimization, pairwise alignment, or a fixed number of reference images, limiting their flexibility and practical use. We introduce WEAVE, a feed-forward framework that supports variable numbers of references without per-scene fine-tuning. WEAVE estimates camera poses and depths in a single forward pass and uses them to reproject each reference into the target view. These projections form two complementary conditioning inputs: a merged query warp (Q-Warp) providing a target-aligned estimate, and separate reference warps (R-Warps) retaining view-specific details and disagreements that merging would discard. WEAVE combines these inputs with the original references to resolve cross-view discrepancies and synthesize the missing content. Experiments on MegaDepth and RealBench show consistent improvements with both MM-DiT and U-Net backbones. WEAVE achieves state-of-the-art reconstruction performance on both benchmarks, outperforming scene-finetuned methods on RealBench without test-time optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.