VAR: Mining Visual Anchors from Vision Encoders for Entity Retrieval in Knowledge-Based Visual Question Answering
Abstract
In Knowledge-Based Visual Question Answering (KB-VQA), correct answer generation relies on the evidence contained in the retrieved gold document. Therefore, accurately retrieving the gold document is critical. This requires matching the question to entities in the knowledge base. However, standard vision encoders often underemphasize subtle, entity-discriminative cues, entangling them with irrelevant background noise in the final global representation. While recent methods address this via query-side region cropping or knowledge-base-side captioning, they either rely on brittle explicit region selection or incur substantial preprocessing overhead. Moreover, existing verification strategies often rely on binary relevance decisions, discarding the relative strength of visual support among candidates. In this paper, we reveal an overlooked phenomenon: fine-grained, entity-discriminative semantics already persist in high-level token representations, merely underexpressed during global aggregation. We conceptualize these compact, retrieval-critical token subsets as Visual Anchors. Based on this insight, we propose VAR (Visual Anchors for Entity Retrieval), a fine-grained entity retrieval framework. VAR first mines visual anchors and uses them to guide [CLS] updates through the remaining frozen encoder layers, strengthening their contribution to the retrieval representation without pixel-space region extraction or knowledge-base index reconstruction. To further improve gold-document ranking, VAR leverages concise visual-attribute descriptions from top-ranked candidates. Beyond binary verification, VAR uses a VLM calibrator to estimate continuous visual support and trains a lightweight gate to jointly calibrate the relative visual support of a candidate set against their retrieval confidence for Top-1 document correction. Extensive experiments demonstrate that VAR achieves state-of-the-art performance on E-VQA, InfoSeek, and OK-VQA, showing the effectiveness and efficiency of our method. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.