RadDiff: Automated Radiology Discovery from Cohort Differences
Abstract
Many clinical questions in radiology are questions about cohorts: Which radiographic findings distinguish pneumonia patients who die from those who survive? How does COVID-19 present in older versus younger patients? What separates the radiographs a race classifier assigns to different groups? Answering them requires discovering, in open-ended language, what distinguishes two cohorts of hundreds of studies, a task that is laborious for radiologists and on which general-domain methods fail. We introduce RadDiff, a multimodal agent for automated radiology discovery from cohort differences. RadDiff runs a hypothesize-verify-refine loop: a proposer reads domain-adapted captions and images of a few sampled studies and proposes candidate differences, a ranker verifies every candidate on the full cohorts with a medical vision-language model, and the best-verified hypotheses, with zoomed-in views of the regions they refer to, are fed back to refine the next round. On RadDiffBench, a new benchmark of 57 radiologist-validated cohort pairs from MIMIC-CXR, RadDiff achieves 47.4% top-1 accuracy versus 1.8% for the general-domain VisDiff, with every component contributing. Applied to the three questions above, it surfaces radiologist-reviewed discoveries, from the device burden of pneumonia non-survivors to the non-anatomical cues behind a race classifier's predictions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.