acceptodds
Under review as a conference paper at ICLR 2027

CAPE: A Clinical Chart Agent for Precision Vision Extraction

Abstract

Multimodal large language models (MLLMs) can answer questions about charts yet struggle to extract their quantitative content precisely and comprehensively. The difficulty is particularly pronounced in clinical charts, where specialized measurements appear within dense visual layouts. Accurate recovery of this information is essential for further quantitative analysis and evidence synthesis. To close this gap, we introduce CAPE, an MLLM-driven agent harness that combines semantic reasoning with pixel-level measurement for clinical chart extraction. The MLLM identifies the quantities to extract and locates the relevant chart elements. Dedicated vision tools follow this guidance to measure the plotted values and return reliability signals that inform remeasurement within a staged workflow. Record-level verification checks the completeness and organization of the extraction before the workflow concludes. To evaluate extraction quality, we develop a novel synthetic data generation approach that produces clinical charts paired with ground truth restricted to quantities readable or derivable from each figure. Using this approach, we construct a benchmark of 950 figures spanning 19 clinical chart types. CAPE raises the Overall extraction score on this benchmark from 0.68 to 0.78 over the strongest baseline built on the Claude Agent SDK and achieves higher scores on 17 of the 19 chart types, producing more complete and correctly organized records. CAPE also achieves the highest Overall score among all evaluated methods on a real clinical chart dataset.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.