acceptodds
Under review as a conference paper at ICLR 2027

ClinLens: Towards Long-Horizon LLM Agents for Longitudinal Multimodal Clinical Data Science

Abstract

Automating clinical data science requires agents to translate patient-centric objectives into executable workflows grounded in longitudinal multimodal evidence. Existing benchmarks provide limited coverage of two essential capabilities: patient-episode grounding and asynchronous multimodal integration. We introduce ClinLens, a patient-centric benchmark for evaluating large language model (LLM) agents on clinical data science. ClinLens comprises 126 executable tasks spanning five data modalities, four analytical scopes, and five analytical capabilities. These tasks assess agents' ability to maintain a coherent patient-specific analytical context across interdependent stages of evidence grounding, multimodal integration, and analytical pipeline generation and execution. We evaluate general-purpose, coding, and biomedical agents. Analyses average 36.93 steps and 8.06 minutes per task. The strongest evaluated agent achieves 86.4% execution success and 54.0% final-answer accuracy, but only a 29.4% strict pass rate. Failure analysis further identifies patient-episode grounding as the most frequently observed failure under the evaluation procedure. By jointly evaluating analytical outcomes and the validity of their supporting clinical context, ClinLens provides a testbed for advancing reliable agents for longitudinal clinical data science.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.