acceptodds
Under review as a conference paper at ICLR 2027

Towards Perception-Centric Multimodal LLM for Multi-Sequence Breast MRI Diagnosis

Abstract

Multimodal large language models (MLLMs) have demonstrated remarkable promise in computer-assisted diagnosis, yet their clinical efficacy remains severely constrained in multi-sequence breast MRI diagnosis. We observe that existing models can achieve better diagnostic conclusions supplied with verified perceptual findings, yet their performance degrades drastically when operating directly on raw multi-sequence MR images, even under explicit clinical guidance. This discrepancy indicates that the primary bottleneck stems from visual perception. We attribute this deficiency to two primary obstacles, specifically clinically unguided perceptual exploration alongside weak multi-sequence perceptual capability. Current MLLMs fail to organize visual interpretation along established clinical perception logic, overlooking subtle yet diagnostically critical findings. Moreover, extracted perceptual cues are weakly grounded across MRI sequences, propagating perceptual errors into downstream diagnostic reasoning. To this end, we propose PCBreast, a perception-centric MLLM for multi-sequence breast MRI diagnosis. We construct MMBreast, a comprehensive multi-center, multi-sequence breast MRI dataset paired with perceptual Chain-of-Thought (CoT) trajectories, enabling the model to internalize clinically aligned perceptual exploration. Furthermore, we introduce Perception-Sensitive Policy Optimization (PSPO), which incorporates perception-grounded rubric reward and visual-sensitive credit assignment to incentivize faithful perceptual grounding and penalize spurious visual findings. Extensive experiments demonstrate that PCBreast consistently outperforms existing MLLMs, achieving impressive improvements in diagnostic accuracy. Our code and model will be publicly available at https://anonymous.4open.science/w/PCBreast-2D2A/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.