V-CUES: Improving Multimodal Search Agents by Centering Visual-Cue Utilization via Experience and Skills
Abstract
Multimodal Search Agents (MSAs) have emerged as a promising paradigm for solving real-world tasks by combining visual understanding with external tools. While recent works have improved the mechanics of tool use and information retrieval, agents still struggle to effectively utilize visual cues during task execution. As a result, agents remain susceptible to Visual-Cue Omission and Visual-Cue Misuse, which can lead to irrelevant retrieval and task failure. To address these limitations, we propose V-CUES, a framework that bootstraps reusable knowledge centered on visual-cue utilization. V-CUES extracts recurring patterns from historical trajectories to capture how visual cues should be utilized under different input contexts, while dynamically refining them through case-level experiences accumulated from new interactions. During execution, V-CUES contextualizes accumulated knowledge for each newly encountered task, providing task-specific supervision on how to incorporate visual cues throughout the search process. Evaluation across four knowledge-intensive benchmarks and three basic models shows that V-CUES consistently improves performance over the previous state-of-the-art methods, demonstrating the effectiveness of explicitly centering visual cues for multimodal search agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.