VERSE: Verifiable Spatial Self-Evolution with Geometric Process Rewards
Abstract
Spatial reasoning remains a fundamental bottleneck for vision–language models (VLMs), largely because models reason purely in 1D text tokens without external geometric anchors, whereas humans routinely sketch and mark up scenes. Equipping VLMs with agentic visual tools narrows this gap, yet training such tools traditionally demands prohibitively expensive human-annotated reasoning trajectories. While the self-evolution paradigm eliminates human annotation in code and mathematics, directly porting it to tool-assisted spatial reasoning encounters a critical failure mode: outcome-only rewards are completely blind to intermediate tool fidelity. Consequently, a policy can draw arbitrary annotations or cross physical obstacles, yet still be falsely reinforced via lucky textual guesses. To address this, we propose VERSE, an annotation-free framework that trains a VLM to actively propose, draw, and self-evolve within 3D environments. At the core of VERSE is the Agentic Verifier: by exploiting underlying 3D scene metadata and procedural topologies, it evaluates intermediate visual tool calls in closed form via computing deterministic grounding IoUs for bounding boxes and sequence-alignment fidelities for traced routes. This yields dense, engine-verified process rewards without relying on learned reward models or LLM judges. Extensive experiments across video, multi-view, and navigation domains demonstrate substantial improvements across scales: on Qwen2.5-VL-7B, VERSE elevates VSI-Bench accuracy from 33.68% to 40.52% and boosts MAZE performance from 33.50% to 90.80%, confirming that engine-verified process rewards provide the critical supervisory signal for multimodal self-improvement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.