DiVE: Test-Time Scaling of Direct Visual Exploration via Cross-Trajectory Verification
Abstract
Knowledge-based visual question answering (KB-VQA) requires identifying the entity depicted in an image and retrieving facts about that entity. However, visually similar long-tail entities can confound grounding. We observe a failure pattern, termed a grounding trap: an agent may retrieve genuine facts about a visually similar but incorrect entity, yielding an internally coherent yet incorrect trajectory. To address this failure, we propose DiVE, a test-time scaling framework that enables direct visual exploration. DiVE combines fine-grained evidence access with direct interaction tools. At inference time, independent rollouts explore alternative entity hypotheses. An off-the-shelf verifier compares the recorded candidate entities and supporting evidence across these trajectories to select an existing answer. Compared with RAG, agentic-search, and alternative test-time scaling baselines, DiVE achieves state-of-the-art performance on E-VQA and InfoSeek. On shared rollout pools, CTV achieves higher accuracy with fewer verifier-stage tokens than the TTS baselines, indicating better verifier-stage efficiency. Our code is available anonymously.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.