acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Reconsider: Adequacy-Guided Inference-Time Interventions for Histopathology VLM Agents

Abstract

Visual agentic systems that work with very large images often cannot process the whole image at once. Instead, they inspect selected regions (patches), ask a large vision-language model (LVLM) to describe what is visible, and use these descriptions to guide later navigation and reasoning. Inadequate or visually irrelevant descriptions can therefore affect the performance of the downstream task. Existing correction methods applied to LVLMs during test-time typically modify attention, decoding, or prompts to reduce visually unsupported responses, but each leaves the agentic system with a decision that cannot be resolved simply by using a stronger model: when should a frozen LVLM perceptor be corrected, and which correction should be used? We study this question in computational pathology, where agents navigate gigapixel scans of tissue specimens known as whole-slide images (WSIs) for visual question answering (VQA), using four LVLMs, three interventions, and two whole-slide image agents. Our experiments reveal two regularities. Descriptions judged relevant, sufficiently complete, and consistent with the ground truth evidence -which we call "adequate" descriptions-support higher downstream answer accuracy than descriptions judged inadequate. Secondly, intervention benefits depend on the model, question type, and baseline evidence quality of the initial description, and the same intervention can repair some descriptions while damaging others. To test whether these empirical patterns transfer to WSI VQA agentic systems and can support useful decisions by selecting one intervention family for low-confidence descriptions and accepting candidates only when their estimated adequacy increases by a fixed margin, we propose RESCUE-ROUTE-RESCORE, a simple controller. On a held-out patch-level benchmark, the proposed controller raises macro description adequacy from % to % and answer accuracy from % to %. Transferred to agentic systems that perform WSI level VQA without any WSI-level tuning, it achieves the highest point estimate of macro answer accuracy in all six dataset-agent comparisons. These findings suggest that intermediate evidence quality can serve as a transferable control signal and that visual agentic systems can improve by learning which evidence to reconsider. Code will be made available on acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.