acceptodds
Under review as a conference paper at ICLR 2027

RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

Abstract

Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often compromised by frequent errors in visual interpretation. To systematically trace these failures, we traverse the hierarchy from high-level clinical tasks down to fundamental visual perception. We therefore introduce Perception-Bench, a large-scale benchmark that assesses medical MLLMs across six dimensions: attribute judgment, spatial grounding, spatial understanding, disease prediction, anomaly detection, and report generation, spanning both 2D and 3D radiology images. Our analysis on Perception-Bench reveals that existing MLLMs lack the ability to capture even the most basic lesion attributes, such as location, size, and density. This failure to ground clinical predictions in visual evidence reveals a critical yet overlooked source of diagnostic unreliability: deficient low-level visual perception. Motivated by this, we propose RadSight, a perception-driven MLLM with dedicated 2D and 3D visual encoders that preserve native imaging spatial structures. RadSight is trained on an 8.37-million-sample perception-oriented corpus through a four-stage curriculum: visual-language alignment, fine-grained visual perception, clinical diagnosis, and diagnostic interpretation. This curriculum explicitly develops perceptual capabilities before progressing to diagnosis and report generation. RadSight outperforms existing MLLMs across all six dimensions of Perception-Bench, with particularly strong gains in spatial grounding and clinical diagnosis. It also achieves consistent improvements on public 2D and 3D medical benchmarks, further demonstrating that robust low-level visual perception is a critical foundation for reliable clinical understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.