Super-Generalist: Generalist–Specialist Synergy for Accurate 3D Computed Tomography Diagnosis and Lesion Grounding across Diverse Diseases
Abstract
Vision–language models offer a promising generalist approach to medical diagnosis, providing broad disease coverage and zero-shot diagnostic capabilities. In 3D computed tomography (CT), however, these generalist models often lack fine-grained anatomical and lesion awareness for reliable diagnosis and spatial interpretability. In contrast, supervised specialist models achieve strong performance on specific tasks but typically lack generalization across diseases and anatomies. In this work, we present SuG, a Super-Generalist framework for 3D CT diagnosis that unifies generalist vision–language learning with specialist objectives, enabling zero-shot multi-disease diagnosis and lesion grounding alongside specialist-competitive diagnostic performance. We perform specialist-enhanced vision-language alignment in SuG by incorporating spatial priors from multiple segmentation experts, including anatomy, class-specific lesion and class-agnostic lesion segmentors that capture lesions beyond anatomies annotated during training. To improve lesion grounding capability, we leverage lesion masks as spatial priors to calibrate text-conditioned visual attention, encouraging disease-related semantics to focus on clinically relevant regions. We evaluate SuG on extensive chest and abdominal CT benchmarks, including CT-RATE, Merlin, MedVL-CT69K, and several in-house tumor datasets. SuG achieves state-of-the-art performance across a wide range of disease diagnosis tasks and surpasses specialist models on several critical tumor diagnosis benchmarks. Furthermore, SuG demonstrates strong lesion grounding capability, including robust generalization to lesion types lacking class-specific supervision. Code and models will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.