acceptodds
Under review as a conference paper at ICLR 2027

From Image Space to Patient Space: Learning Patient-Centered Coordinates in Medical MLLMs

Abstract

Medical MLLMs are increasingly moving beyond visual recognition toward more complex medical reasoning, yet they still struggle with spatial-relation reasoning in medical images. Our analyses reveal that medical MLLMs already encode image-space spatial information in the intermediate representations. However, this image-space representation alone is insufficient for clinical spatial reasoning. We therefore propose *Patient-Space Coordinate Learning (PS-Coord)*, a dual-frame training framework that jointly supervises image-space and patient-space coordinates within intermediate hidden states. Across eight backbones spanning 7B–34B parameters, PS-Coord consistently improves CT spatial reasoning, with average gains of 19.67% and 16.17% on MIRP and MSD, respectively. The improvements further generalize to unseen imaging modalities, held-out rotations, and medical visual question answering. Representation analyses show substantially increased accessibility of patient coordinates after training, while counterfactual controls indicate that PS-Coord-trained models rely on the anatomical context of the current image for medical spatial decisions. These results suggest that explicitly connecting image and patient coordinate provides an effective way for enabling spatial reasoning capabilities already latent within medical MLLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.