acceptodds
Under review as a conference paper at ICLR 2027

Beyond Time: A Physically-Grounded Two-Stage VLM for 3D Medical Image Segmentation via Spatial Prompting

Abstract

Bringing temporal modeling from vision-language models into 3D medical image understanding offers a practical way to process full-length CT and MRI volumes and improve training stability. Existing methods still have two main limits. They often treat the spatial Z axis as a uniformly sampled time axis and ignore the true physical slice thickness, which weakens spatial perception. They also rely on external foundation models for one-way slice-by-slice mask propagation, which can accumulate errors in long sequences. We propose PhysiVLM, a physically constrained two-stage framework that explicitly separates 3D spatial reasoning from 2D mask generation. In the first stage, we use an anatomical 2.5D grid puzzle and spatial coordinate prompts to inject the prior of physical slice thickness, which allows the model to regress the global 3D depth range of a lesion at low cost and with high accuracy. In the second stage, we introduce a Spatial Prompting mechanism to build anchor-guided reasoning. The model takes the predicted depth range from the first stage, dynamically samples key slices, and then guides the VLM to perform slice-wise spatial semantic reasoning, parse lesion shape, and output structured 2D bounding boxes. We then feed these boxes into MedSAM2 to produce spatially consistent bidirectional mask propagation along the Z axis. Extensive experiments show that PhysiVLM reduces cumulative errors in long-sequence propagation, cuts redundant computation, and improves segmentation accuracy for lesions with complex cross-slice morphology. The code is available at: https://anonymous.4open.science/r/PhysiVLM-1E33

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.