DRIVE: Dual Reflection with Interleaved Visual Evidence for Embodied Pointing
Abstract
A model may identify the intended visual target yet predict a coordinate that falls elsewhere. We present DRIVE (Dual Reflection with Interleaved Visual Evidence), a two-stage reinforcement-learning framework designed to strengthen the connection between spatial reasoning and coordinate prediction. A shared policy first proposes target points, then reflects on their placement through rendered feedback: it states a spatial expectation before rendering and diagnoses the observed marker afterward. When the diagnosis identifies a mismatch, the policy revises the coordinates and inspects them again. Training uses final localization outcomes to optimize a selected response conditioned on this reflection history. On the 300-example VABench-P test set, the complete protocol achieves 71.67% final-hit accuracy, compared with 67.33% for direct inference with the same trained checkpoint. Inference ablations examine how spatial expectations, placement diagnoses, and rendered markers contribute to the trained policy’s predictions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.