Towards Multi-CT Perception and Understanding with a Radiology Tool-Augmented Multimodal Large Language Model
Abstract
Clinical CT interpretation inherently requires multi-volume reasoning, as radiologists routinely synthesize complementary evidence across anatomical regions, contrast phases, and longitudinal visits for accurate diagnostic assessment. However, existing 3D medical multimodal large language models predominantly operate on a single CT volume and fail to integrate complementary evidence across scans and longitudinal visits that radiologists routinely consider in clinical decision-making. Meanwhile, large-scale benchmarks for evaluating such multi-volume CT reasoning remain limited. To address these limitations, we formulate a challenging yet important setting, termed multi-CT perception and understanding, in which a model jointly reasons over multiple CT volumes. First, we construct a large-scale benchmark, namely CliMCT-Bench from 87,953 clinical visits involving 298,377 CT volumes, resulting in 267K visual question-answer pairs spanning six clinical capabilities from basic perception to longitudinal reasoning. We further propose -FM, a multimodal large language model that integrates evidence across scans and interleaves reasoning with the adaptive invocation of radiology tools. Beyond standard contrastive pretraining and supervised fine-tuning, we introduce two innovative training strategies. Specifically, we propose multi-turn and multi-scale trajectories to enable the model to perceive both global context and salient regions, facilitating coarse-to-fine visual reasoning. Meanwhile, tool-augmented trajectories interleave reasoning with radiology tool invocation, encouraging the model to adaptively acquire additional visual evidence. Extensive experiments show that our method significantly outperforms existing state-of-the-art models by 12.41% in average accuracy, while the ablation and user studies demonstrate the effectiveness of tool invocation. Code and benchmark will be available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.