CAPE: A Clinical Chart Agent for Precision Vision Extraction
Abstract
Multimodal large language models (MLLMs) can answer questions about charts yet struggle to extract their quantitative content precisely and comprehensively. The difficulty is particularly pronounced in clinical charts, where specialized measurements appear within dense visual layouts. Accurate recovery of this information is essential for further quantitative analysis and evidence synthesis. To close this gap, we introduce CAPE, an MLLM-driven agent harness that combines semantic reasoning with pixel-level measurement for clinical chart extraction. The MLLM identifies the quantities to extract and locates the relevant chart elements. Dedicated vision tools follow this guidance to measure the plotted values and return reliability signals that inform remeasurement within a staged workflow. Record-level verification checks the completeness and organization of the extraction before the workflow concludes. To evaluate extraction quality, we develop a novel synthetic data generation approach that produces clinical charts paired with ground truth restricted to quantities readable or derivable from each figure. Using this approach, we construct a benchmark of 950 figures spanning 19 clinical chart types. CAPE raises the Overall extraction score on this benchmark from 0.68 to 0.78 over the strongest baseline built on the Claude Agent SDK and achieves higher scores on 17 of the 19 chart types, producing more complete and correctly organized records. CAPE also achieves the highest Overall score among all evaluated methods on a real clinical chart dataset.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.