acceptodds
Under review as a conference paper at ICLR 2027

AnatoBridge: Strengthening and Transferring Anatomical Representations for 3D Chest CT Vision-Language Models

Abstract

Multimodal large language models remain challenged in understanding 3D chest computed tomography (CT), where clinically relevant evidence is volumetric, spatially localized, and often subtle. Existing models inadequately exploit the anatomical structure needed to ground such evidence: i) anatomical guidance is weakly integrated into visual encoding, ii) lesion supervision remains underused in representation learning, and iii) explicit anatomical outputs are not directly exposed to the LLM. In this paper, we present AnatoBridge, an anatomy-aware 3D chest CT VLM that incorporates anatomical information throughout visual encoding, representation learning, and language alignment to address these limitations. AnatoBridge strengthens the visual representations by integrating region-constrained anatomy queries throughout the transformer and using lesion-guided pretraining to enrich organ features and construct directly supervised disease-specific representations. Moreover, a dual-channel interface carries this information into language modeling to provide the LLM with complementary fine-grained visual information and interpretable anatomical evidence, where the feature channel projects the learned representations into continuous anatomy tokens and the prompt channel serializes explicit anatomical outputs as disease findings, organ geometry, and lesion grounding. Experimental results show that AnatoBridge outperforms the strongest baseline by 1.1 in AUC (85.5 vs. 84.4) in zero-shot classification on CT-RATE and yields a 0.7 AUC gain (76.1 vs. 75.4) on the external Rad-ChestCT cohort. Remarkably, it also exceeds RadSight-8B by an average 3.33 points across the three 3D-RAD tasks, by 1.80 points on CT-Spatial-VQA, and by 6.23 points in report-generation clinical efficacy F1.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.