ClinLens: Towards Long-Horizon LLM Agents for Longitudinal Multimodal Clinical Data Science
Abstract
Automating clinical data science requires agents to translate patient-centric objectives into executable workflows grounded in longitudinal multimodal evidence. Existing benchmarks provide limited coverage of two essential capabilities: patient-episode grounding and asynchronous multimodal integration. We introduce ClinLens, a patient-centric benchmark for evaluating large language model (LLM) agents on clinical data science. ClinLens comprises 126 executable tasks spanning five data modalities, four analytical scopes, and five analytical capabilities. These tasks assess agents' ability to maintain a coherent patient-specific analytical context across interdependent stages of evidence grounding, multimodal integration, and analytical pipeline generation and execution. We evaluate general-purpose, coding, and biomedical agents. Analyses average 36.93 steps and 8.06 minutes per task. The strongest evaluated agent achieves 86.4% execution success and 54.0% final-answer accuracy, but only a 29.4% strict pass rate. Failure analysis further identifies patient-episode grounding as the most frequently observed failure under the evaluation procedure. By jointly evaluating analytical outcomes and the validity of their supporting clinical context, ClinLens provides a testbed for advancing reliable agents for longitudinal clinical data science.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.