Med3DNavi: Agentic Visual Exploration of 3D CT Volumes via a Structured Volumetric Reading Protocol
Abstract
Three-dimensional computed tomography (CT) is central to diagnosing abdominal, thoracic, and neurological conditions, yet vision-language models (VLMs) natively perceive only 2D images. Prior work bridges this gap in two ways, each with a characteristic failure mode. Fixed 2D conversions, whether uniformly sampled slices or a trained 3D encoder, commit what the model sees before it knows what it needs, yielding **vision without agency**. Text-based medical agents instead reason over tool reports without ever inspecting the pixels, yielding **agency without vision**. We argue that both limitations trace to a single root cause: the VLM is never placed in a closed-loop visual dialogue with the volume, unlike a radiologist who actively navigates and cross-checks across planes and windows before concluding. Consequently, we present **Med3DNavi**, a training-free agent that couples a programmable CT viewer, which renders on-demand multi-planar slices, grids, HU windowing, ROI zoom, and segmentation-mask overlays, with a Structured Volumetric Reading Protocol (SVRP). SVRP follows an orient–investigate–adjudicate workflow, routing each case either to an audit path for verifying an available tool prior or to an autonomous path for proactive and independent visual exploration. Experiments on DeepTumorVQA and 3D-RAD show that Med3DNavi improves over competitive baselines, with benefits that generalize across backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.