One Lesion, Many Views: Disentangled Vision-Language Pretraining for Multi-Dimensional Lesion Understanding
Abstract
Vision-Language Pretraining (VLP) has emerged as a powerful paradigm for medical image understanding by leveraging unstructured reports. However, existing medical VLP methods often fall short in fine-grained lesion understanding, primarily due to the inherent sparsity of lesion-specific information within free-text reports and the lack of explicit lesion-level anchoring during alignment. To address these limitations, we propose Lesion-CLIP, a disentangled VLP framework for multi-dimensional lesion understanding. First, for what to align, we introduce multi-dimensional data decoupling by decomposing 3D CT scans and radiology reports into five distinct semantic dimensions. Second, for where to align, we develop dimension-specific space decoupling by projecting image features into distinct subspaces. Third, for how to align, we design asymmetric direction decoupling to perform hierarchical dimension-level and lesion-level alignment. To this end, we curate LiLA, a pathology-calibrated 3D CT lesion dataset spanning type, size, location, phase, and attribute to facilitate fine-grained training and evaluation. Extensive experiments on LiLA and the public MSWAL dataset demonstrate that Lesion-CLIP achieves state-of-the-art performance across both discriminative and generative tasks. Notably, zero-shot Lesion-CLIP consistently outperforms established fully supervised models across all core dimensions on LiLA, achieving an absolute accuracy gain of 7.56% in lesion classification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.