Beyond Time: A Physically-Grounded Two-Stage VLM for 3D Medical Image Segmentation via Spatial Prompting
Abstract
Bringing temporal modeling from vision-language models into 3D medical image understanding offers a practical way to process full-length CT and MRI volumes and improve training stability. Existing methods still have two main limits. They often treat the spatial Z axis as a uniformly sampled time axis and ignore the true physical slice thickness, which weakens spatial perception. They also rely on external foundation models for one-way slice-by-slice mask propagation, which can accumulate errors in long sequences. We propose PhysiVLM, a physically constrained two-stage framework that explicitly separates 3D spatial reasoning from 2D mask generation. In the first stage, we use an anatomical 2.5D grid puzzle and spatial coordinate prompts to inject the prior of physical slice thickness, which allows the model to regress the global 3D depth range of a lesion at low cost and with high accuracy. In the second stage, we introduce a Spatial Prompting mechanism to build anchor-guided reasoning. The model takes the predicted depth range from the first stage, dynamically samples key slices, and then guides the VLM to perform slice-wise spatial semantic reasoning, parse lesion shape, and output structured 2D bounding boxes. We then feed these boxes into MedSAM2 to produce spatially consistent bidirectional mask propagation along the Z axis. Extensive experiments show that PhysiVLM reduces cumulative errors in long-sequence propagation, cuts redundant computation, and improves segmentation accuracy for lesions with complex cross-slice morphology. The code is available at: https://anonymous.4open.science/r/PhysiVLM-1E33
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.