acceptodds
Under review as a conference paper at ICLR 2027

Med3DNavi: Agentic Visual Exploration of 3D CT Volumes via a Structured Volumetric Reading Protocol

Abstract

Three-dimensional computed tomography (CT) is central to diagnosing abdominal, thoracic, and neurological conditions, yet vision-language models (VLMs) natively perceive only 2D images. Prior work bridges this gap in two ways, each with a characteristic failure mode. Fixed 2D conversions, whether uniformly sampled slices or a trained 3D encoder, commit what the model sees before it knows what it needs, yielding **vision without agency**. Text-based medical agents instead reason over tool reports without ever inspecting the pixels, yielding **agency without vision**. We argue that both limitations trace to a single root cause: the VLM is never placed in a closed-loop visual dialogue with the volume, unlike a radiologist who actively navigates and cross-checks across planes and windows before concluding. Consequently, we present **Med3DNavi**, a training-free agent that couples a programmable CT viewer, which renders on-demand multi-planar slices, grids, HU windowing, ROI zoom, and segmentation-mask overlays, with a Structured Volumetric Reading Protocol (SVRP). SVRP follows an orient–investigate–adjudicate workflow, routing each case either to an audit path for verifying an available tool prior or to an autonomous path for proactive and independent visual exploration. Experiments on DeepTumorVQA and 3D-RAD show that Med3DNavi improves over competitive baselines, with benefits that generalize across backbones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.