Framing Abnormalities in Context: Reciprocal Grounding for Radiograph Vision–Language Pre-training
Abstract
Medical vision–language pre-training benefits from the rich supervision contained in radiology reports. However, existing methods often model focal clinical findings and their surrounding descriptions independently, without explicitly coordinating the complementary semantics they provide. We propose FRAME (Finding–context Reciprocal Alignment with Masked report reconstruction and Evidence-focused summary), a radiograph pre-training framework that couples finding-level and sentence-level semantics through shared visual evidence. Specifically, finding-conditioned visual representations are aligned with their corresponding sentence contexts, while context-conditioned visual representations are reciprocally aligned with the associated findings. This bidirectional constraint encourages the visual encoder to preserve both abnormality-specific evidence and its broader clinical context. Two complementary proxy tasks further strengthen the learned representation. Masked report reconstruction promotes the retention of dense diagnostic information, whereas evidence-focused summary emphasizes concise and clinically salient semantics. We evaluate FRAME across classification, segmentation, detection, visual question answering, and report generation under limited-label and cross-dataset settings. FRAME consistently outperforms strong Med-VLP baselines across these task families, especially with only 1% of the training labels. Ablation studies further demonstrate the advantage of reciprocal over unidirectional alignment and verify the complementary contributions of both proxy objectives. These results establish reciprocal finding–context grounding as an effective mechanism for learning transferable radiographic representations. Code will be available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.