Training-Free Looped Vision Geometry Transformers
Abstract
Can the geometric predictions of pretrained feed-forward models be improved without retraining or fine-tuning? We study training-free looped inference across VGGT, Pi3, Depth Anything 3, MapAnything, and VGGT-. An architecture-aware wrapper repeats a contiguous block window with damped updates, preserving pretrained weights, native attention, and prediction heads. We compare original (raw) and loop-augmented (loop) inference using one fixed configuration per model across static and dynamic benchmarks. Looping raises VGGT's HiRoom reconstruction F1 from 53.79 to 60.43; MapAnything improves every measured reconstruction cell across five benchmarks. Benefits depend on the task and window. On HiRoom, higher raw patch-feature similarity within a window is associated with larger F1 gains for VGGT and Pi3 after accounting for window length. Consistently updating intermediate features also enables useful loops across layers supplying dense prediction heads. These results demonstrate training-free improvements to multi-view geometry and connect their benefits to the structure of pretrained representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.