acceptodds
Under review as a conference paper at ICLR 2027

ClinSeekAgent: Automating Multi-modal Evidence Seeking for Agentic Clinical Reasoning

Abstract

Large language models (LLMs) and agentic systems have shown promise for clinical decision support, but existing works largely assume that evidence has already been curated and handed to the model. Real-world clinical workflows instead require agents to actively seek, iteratively plan, and synthesize multimodal evidence from heterogeneous sources. In this paper, we introduce **ClinSeekAgent**, an automated agentic framework for dynamic multimodal evidence seeking that shifts the paradigm from passive evidence consumption to active evidence acquisition. Given a clinical task and access to raw patient data, ClinSeekAgent queries longitudinal EHR records and invokes imaging analysis tools when relevant; refines its hypotheses as new information emerges; and integrates the collected evidence into grounded clinical decisions. We validate ClinSeekAgent in two roles: as an inference-time agent for frontier LLMs and as a training pipeline for open-source models. For inference-time validation, we construct **ClinSeek-Bench**, which covers 13 text-only EHR task categories and 6 multimodal EHR–chest X-ray task categories, and pairs each question with both a *Curated Input* of pre-selected evidence and the raw patient data for automated evidence seeking. Without any benchmark-specific curation, ClinSeekAgent significantly outperforms curated input for both Claude Opus 5 and GPT-5.6 Sol, improving overall F1 from 61.7 to 63.9 and from 57.8 to 59.3 on text-only EHR tasks, and from 66.5 to 68.8 and from 60.0 to 62.5 on multimodal tasks. The largest gains occur in text-based decision-making and multimodal phenotype prediction, and ablations show that they come from multi-round evidence seeking rather than from access to more of the patient record. For training-time validation, we fine-tune Qwen3.5-35B-A3B on ClinSeekAgent trajectories generated by GPT-5.6 Sol. The resulting **ClinSeek-35B-A3B** reaches 35.6 average F1 on AgentEHR-Bench (**+16.1** over its base model), outperforming all evaluated open-source models, including ones more than ten times larger, as well as Claude Opus 5. We will release our model, data, and code to facilitate future research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.