CT-AGENT: AN EVIDENCE-GROUNDED MULTI-TOOL AGENT FOR HETEROGENEOUS CT ANALYSIS
Abstract
Clinical CT interpretation comprises a sequence of evidence-gathering steps. A radiologist moves across slices, localizes findings, measures lesions, compares examinations, and reconciles these observations before producing a report. Current 3D vision-language models usually compress this process into a label, answer, or generated report. Specialist imaging models expose useful intermediate results and usually operate as separate tools. We present CT-Agent, an evidence-grounded agent that organizes image-only, text-only, and image-text requests as structured tasks, invokes expert models through the Model Context Protocol, and retains masks, bounding boxes, measurements, text records, and execution provenance in a case-level evidence memory. The same evidence supports quantitative reasoning and output verification. Each capability is evaluated with a task-specific benchmark. On LiTS, CT-Agent improves mean Dice from 0.715 to 0.731 and reduces measurement errors relative to 3DMedAgent. On 40 BronAtlas cases, coordinating ATM22 and AeroPath raises Dice to 0.750 and precision to 0.993. For abdominal screening, organ-level inputs, a lightweight adapter, and calibrated thresholds recover the zero sensitivity observed under direct CT-CLIP transfer. In a 40-case CT-RATE workflow, every case completes tool execution and evidence export; verification corrects four report assertions and introduces one error. These results characterize the effects of tool orchestration and the current limits of calibration, sample size, and automatic verification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.