acceptodds
Under review as a conference paper at ICLR 2027

One Lesion, Many Views: Disentangled Vision-Language Pretraining for Multi-Dimensional Lesion Understanding

Abstract

Vision-Language Pretraining (VLP) has emerged as a powerful paradigm for medical image understanding by leveraging unstructured reports. However, existing medical VLP methods often fall short in fine-grained lesion understanding, primarily due to the inherent sparsity of lesion-specific information within free-text reports and the lack of explicit lesion-level anchoring during alignment. To address these limitations, we propose Lesion-CLIP, a disentangled VLP framework for multi-dimensional lesion understanding. First, for what to align, we introduce multi-dimensional data decoupling by decomposing 3D CT scans and radiology reports into five distinct semantic dimensions. Second, for where to align, we develop dimension-specific space decoupling by projecting image features into distinct subspaces. Third, for how to align, we design asymmetric direction decoupling to perform hierarchical dimension-level and lesion-level alignment. To this end, we curate LiLA, a pathology-calibrated 3D CT lesion dataset spanning type, size, location, phase, and attribute to facilitate fine-grained training and evaluation. Extensive experiments on LiLA and the public MSWAL dataset demonstrate that Lesion-CLIP achieves state-of-the-art performance across both discriminative and generative tasks. Notably, zero-shot Lesion-CLIP consistently outperforms established fully supervised models across all core dimensions on LiLA, achieving an absolute accuracy gain of 7.56% in lesion classification.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.