acceptodds
Under review as a conference paper at ICLR 2027

Cardiomni: An Evidence-Grounded Agentic Workflow for Multi-View Multi-Dimension Cardiac MRI Interpretation

Abstract

Cardiac magnetic resonance (CMR) interpretation spans a continuum from low-level visual perception to high-level clinical reasoning: a physician must inspect multi-view cine and LGE sequences, identify abnormal visual patterns, measure ventricular structure and function, and integrate these observations into a diagnosis. Existing CMR analysis systems are usually either task-specific predictors or end-to-end vision-language models that produce one-hop answers, making it difficult to accumulate explicit perceptual and quantitative evidence for downstream reasoning. Meanwhile, general and medical multimodal large language models have improved image-text understanding, but they are not designed to actively acquire CMR-specific evidence such as attribution maps, segmentation-derived volumes, ejection fraction, and indexed measurements. To bridge this gap, we propose Cardiomni, an evidence-grounded agentic workflow for multi-view, multi-dimensional CMR interpretation. Cardiomni coordinates two complementary evidence pathways through an iterative Planner–Executor–Verifier loop: a Visual Attention Highlighter provides abnormality probabilities with class-conditioned spatial evidence, while a Quantitative Measurement Estimator derives clinically interpretable measurements from cardiac masks and scan metadata. A persistent evidence ledger stores intermediate findings, enabling the planner to adaptively decide what evidence is still missing and allowing a frozen verifier and answer generator to produce answers grounded in the collected evidence. We train only the planner with reinforcement learning using a rule-based final-answer-correctness reward, while keeping the visual, quantitative, verification, and generation components fixed. We further introduce DeepCMRVQA, a benchmark for evaluating CMR perception-to-reasoning across 35 tasks covering major abnormality assessment and quantitative measurement interpretation. Experiments on binary abnormality VQA and measurement-range MCQ tasks show that Cardiomni outperforms general-purpose and medical vision-language baselines in balanced accuracy, highlighting a scalable path toward auditable CMR decision-support agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.