acceptodds
Under review as a conference paper at ICLR 2027

Diagno-SV: Decoupling Perception, Reasoning, and Knowledge for Trustworthy Medical VLMs

Abstract

While Large Vision-Language Models (VLMs) show immense promise in emulating the complex cognitive processes of clinical diagnostics, ensuring the reliability of their multimodal reasoning remains a critical and unsolved challenge. Current Chain-of-Thought approaches in VLMs typically conflate visual perception, logical deduction, and medical knowledge, creating an opaque "black box." This ambiguity fosters a critical failure mode we identify as retrospective forgery, a phenomenon where models fabricate visual findings or distort facts post-hoc to justify flawed diagnostic hypotheses. To address this, we introduce Diagno-SV (Diagnosis-Separation for medical VLM), a novel framework for diagnostic-oriented optimization via trilevel decoupling. Diagnosis-SV forces the model to decompose its reasoning into independently verifiable steps of visual perception, logical deduction, and medical knowledge. Building on this transparency, we design a curriculum-based diagnostic refinement workflow. A dedicated multimodal diagnostic module pinpoints errors, distinguishing between "looking wrong" and "thinking wrong", and constructs preference pairs for Direct Preference Optimization. The resulting Diagno-SV model achieves state-of-the-art performance across six medical VQA benchmarks, including MMMU-Med and PMC-VQA. More importantly, our method significantly mitigates retrospective forgery, reducing the rate of visual and factual fabrication from 42.74% to 26.32%, paving a new path toward trustworthy visual-based medical AI.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.