Longitudinal Radiology Report Generation via Progression Alignment and Pathology Semantic Learning
Abstract
Longitudinal radiology report generation requires representations that capture both clinical findings and their changes across examinations. To address these complementary requirements, we propose a unified framework for cross-modal progression alignment and pathology semantic enhancement. We represent progression as the residual between current and prior representations within each modality. Through temporal direction alignment, we encourage consistent directions of change between visual and textual representations, while auxiliary static alignment maintains image–report consistency within each examination. Cross-modal consistency alone, however, does not explicitly associate visual features with clinical findings. To establish these associations, a label-induced inverse optimal transport objective jointly learns pathology prototypes and a visual adapter, providing semantic guidance to the features used for generation. The resulting representations combine current visual evidence with prior image–report context to condition a frozen language model. Experiments on Longitudinal-MIMIC demonstrate strong performance in report generation and clinical agreement, with ablation studies and representation analyses supporting the complementary roles of progression alignment and explicit pathology semantic supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.