acceptodds
Under review as a conference paper at ICLR 2027

Match What You Know: Capability-Aligned Search Agents for Knowledge-Based Visual Question Answering

Abstract

Knowledge-Based Visual Question Answering (KBVQA) requires models to identify task-relevant entities in images and retrieve complementary information from external resources when their internal knowledge is insufficient. Existing retrieval-augmented generation methods rely on predefined retrieval pipelines, which may lead to either insufficient evidence acquisition or redundant retrieval. Although search agents can dynamically select tools, reinforcement learning often drives them towards a conservative full-pipeline retrieval workflow of "image search followed by text search". Consequently, even when the model already possesses the necessary entity-recognition or question-answering capabilities, it still tends to invoke external tools first, leading to underutilization of its existing capabilities and unnecessary tool calls. To address this issue, we propose , a multimodal search agent for KBVQA. Before reinforcement learning, we conduct tool-free behavioral probing to estimate the model’s initial capabilities in three aspects: identifying task-relevant entities, answering questions when the correct entity is provided, and directly answering questions from the original image-question pair. During training, we use the probing results to construct capability-boundary supervision, which encourages the model to prioritize its existing capabilities and invoke external tools only when the corresponding capabilities are insufficient. BoundarySA therefore learns to expand its problem-solving ability through tool use while continuing to leverage its original capabilities. Experiments on MMHops, InfoSeek, and E-VQA demonstrate that BoundarySA maintains competitive answer performance while making more effective use of the model’s existing capabilities and reducing reliance on the fixed workflow.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.