MedInSight: Active Cross-Image Evidence Acquisition for Clinical Diagnosis
Abstract
Clinical diagnosis from heterogeneous medical images requires finding and integrating visual evidence across multiple images. Existing methods lack the ability to actively acquire and align distributed visual evidence, leading to missed subtle visual findings or incorrect cross-image relationships, thereby degrading overall performance. To address these challenges, we propose MedInSight, an end-to-end agentic framework for active cross-image clinical reasoning. Central to MedInSight is Cross-image Evidence Routing (CER), which leverages predictive uncertainty in per-image grounding and spatial uncertainty across images for targeted cross-image inspection. To optimize these visual interactions, we further propose Cross-horizon Credit Assignment (CCA), which combines short-horizon action-level credit magnitude with long-horizon stage-level direction for fine-grained credit assignment beyond uniform outcome-based advantages. Together, CER and CCA couple targeted cross-image evidence acquisition with fine-grained policy optimization, enabling progressive integration of distributed visual evidence for clinical diagnosis. We instill these capabilities through a two-stage training pipeline consisting of supervised fine-tuning (SFT) followed by reinforcement learning (RL). Extensive experiments demonstrate that MedInSight achieves state-of-the-art performance across medical diagnosis and reasoning benchmarks, improving its base model, Qwen3-VL-8B, by 22.96 points on average with consistent gains on out-of-domain datasets. Codes and checkpoints will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.