AnatoRAG: Anatomy-Guided Retrieval for Chest X-ray Report Generation
Abstract
Given radiographic images, chest X-ray report generation aims to produce clinically accurate textual findings. A common recipe is retrieval-augmented generation (RAG), which takes prior reports as reference. However, the reports in such a corpus come from other patients, so their findings are plausible yet unverified for the target image, and report-level retrieval cannot separate the compatible findings from the rest. Yet retrieving at statement level does not resolve this: the retrieved text remains unverified against the target image and largely redundant. We introduce AnatoRAG, which treats retrieved evidence as a hypothesis to be verified and compressed rather than text to be copied. First, AnatoRetriever uses coarse anatomical masks to guide statement retrieval, and a finding classifier discards statements that contradict the target image. Second, the Anatomy Evidence Adapter condenses the surviving statements and image-label predictions into compact anatomical representations, which a multimodal LLM fuses with visual features to generate the report. Extensive experiments on MIMIC-CXR show that AnatoRAG achieves 65.1 CheXbert-14 micro-F1, surpassing RADAR by 2.4 points. Trained only on MIMIC-CXR, AnatoRAG also transfers to the unseen CheXpert Plus and IU X-Ray datasets, yielding 60.4 and 54.6 CheXbert-14 micro-F1, respectively. The improvement is concentrated on findings whose evidence must be recomposed across anatomical regions, indicating that verification and compression, rather than retrieval alone, drive the gain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.