acceptodds
Under review as a conference paper at ICLR 2027

Harnessing Medical Visual-Language Models with Agentic Evidence Management

Abstract

Medical vision–language agents increasingly rely on segmentation tools, and recent work has focused on acquiring the right mask, often treating a good mask as the end point of tool use. However, do these tool-augmented agents actually make effective use of the evidence once it is obtained? To answer this question, we need a benchmark that separates evidence use from segmentation quality, spans multiple medical image types, and allows the same mask to be presented in different ways. Existing medical VQA datasets do not jointly provide these properties. We therefore construct Mask-Evidence MedVQA, a 2.5K-example benchmark spanning four image types, where each question is paired with an answer-relevant ground-truth mask and multiple possible renderings. Our study reveals two limitations of direct tool use. First, evidence presentation matters. Different renderings of the same fixed mask can substantially change model behavior. Second, masks introduce an evidence trade-off. Highlighting a relevant region can expose diagnostic evidence the model previously missed, but the added visual cue can also induce incorrect associations. As a result, even a ground-truth mask can correct some answers while turning initially correct ones into incorrect ones. Together, these findings show that acquiring the right mask is not enough; an agent must also decide how to present the evidence and whether it should change the answer. We therefore introduce TUNE (Tool Use is Not Enough), an evidence-management harness that explicitly manages what evidence to acquire, how to present it, and whether to accept the resulting answer update. Across three standard medical VQA benchmarks, TUNE improves its backbone's mean no-tool accuracy by 11.5 points and exceeds the previous state-of-the-art medical agent by 3.5 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.