Don't Trust, Verify: Multi-Agent Framework with Adversarial Evidence Cross-Examination for Whole Slide Image
Abstract
Whole slide image (WSI) analysis is increasingly approached by agentic systems that emulate a pathologist’s workflow—surveying a slide and zooming into suspicious regions to accumulate evidence. We argue that this exploratory paradigm shares a fundamental blind spot: it over-trusts the perceptions of underlying multimodal large language models (MLLMs) and seeks to improve performance by gathering additional visual observations, even though navigating to different regions often cannot recover features that the model cannot reliably resolve. We introduce ReMIL-Agent, a training-free multi-agent framework operating along an orthogonal verification axis based on multiple instance learning (MIL), in which each MLLM-reported feature is treated as evidence to be actively cross-examined rather than accepted at face value. A cascade mechanism then revisits an initial prediction, subjects its supporting morphological cues to an adversarial Defender–Critic process grounded in a pathology knowledge base, and selectively re-examines key evidence across scales only when warranted, ultimately retreating to UNCERTAIN at the limits of VLM capability. Finally, a reviewer module performs multi-scale reasoning to produce interpretable and auditable pathological reports. Empirical results demonstrate that ReMIL-Agent is a generalizable workflow agent framework that can be built on top of existing MIL methods, enabling more transparent reasoning with explicit capability boundaries and further advancing trustworthiness for clinical pathologists.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.