RRG-MLVA: Radiology Report Generation with Multi-Level Vision–Language Alignment
Abstract
Automated radiology report generation can reduce clinical workload, yet existing methods often struggle with fine-grained vision–language alignment and limited multi-level semantic reasoning. We propose RRG-MLVA, a framework that couples a Swin Transformer encoder and a LoRA-adapted LLaMA-2 decoder with alignment objectives operating at three semantic levels: token-level alignment for fine-grained grounding, multi-scale prototype alignment for hierarchical disease modelling, and disease-level alignment for diagnostic consistency, together with a global contrastive objective. At the core of the framework is a multi-scale Banzhaf formulation that treats learnable disease prototypes as players in a cooperative game and quantifies each prototype's contribution as its expected marginal contribution over sampled coalitions. Because the underlying coalition value function is non-additive, these contributions do not reduce to the individual prototype similarities, and the resulting distributions provide a principled signal for aligning visual and textual evidence. Prototypes are maintained at coarse (7), intermediate (14) and fine (28) granularities, allowing disease semantics to be matched across levels of specificity. Experiments on CheXpert Plus and IU-Xray evaluate the framework with standard language-generation metrics alongside CheXbert and RadGraph clinical-efficacy measures, with ablations isolating the contribution of each objective and of the prototype granularity. Parameter-efficient LoRA fine-tuning keeps the number of trainable parameters small. Results indicate that multi-level alignment improves report accuracy, coherence, and clinical relevance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.