VERVE: Verified Visual Search
Abstract
Recent visual search methods improve perception in high-resolution images, but these gains can come at a cost: reduced ability to recognize when no target satisfies the query (hallucination). We introduce Verified Visual Search (VERVE), a framework for measuring and mitigating this perception-hallucination trade-off. Our VERVE-MCQ benchmark pairs positive visual search queries with minimally edited negative counterparts for which no matching target exists, testing both target identification and query verification. We find that state-of-the-art visual search methods surprisingly underperform their own base models on VERVE-MCQ. To diagnose the root cause, VERVE-Binary simplifies the queries into a yes/no task. Yet the failure persists, indicating a broader loss of discriminativeness. To tackle this, we propose VERVE-Train, a GRPO-based method that trains models to verify rather than assume target existence, adding negative queries (plain mode) and further on-policy self-verification rewards (OPSV mode). Notably, our trained Qwen3.5-9B surpasses Gemini-3.7-Flash by 13.8% on VERVE-MCQ. Crucially, VERVE-Train effectively reduced hallucinations across six benchmarks and remain competitive on eight perception benchmarks. Furthermore, these improvements transfer across tasks, domains, and modalities on seven external benchmarks, demonstrating benefits beyond high-resolution visual search.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.