acceptodds
Under review as a conference paper at ICLR 2027

MedSkill: Internalizing Clinically Grounded Visual Skills for Medical Multimodal Reasoning

Abstract

Medical multimodal large language models (MLLMs) have shown strong performance in medical visual question answering, yet their visual competence is still largely acquired from image–text correlations and task-specific supervision. In clinical practice, however, image interpretation relies on reusable visual procedures for locating anatomy, recognizing findings, assessing spatial relations, and characterizing visual appearance and morphology. These procedures are explicitly reflected in standardized reporting systems and assessment guidelines, but remain largely implicit in current medical MLLMs.We introduce MedSkill, a framework for internalizing clinically grounded visual skills into medical MLLMs. MedSkill derives reusable visual skills from authoritative clinical standards and grounds them in diverse medical images. Through structured supervision, these skills are internalized as persistent visual capabilities, enabling the model to apply clinical visual expertise directly during reasoning.Experiments on multiple medical VQA benchmarks show that MedSkill improves overall question-answering performance and strengthens the corresponding visual skills measured through targeted capability evaluation. Further analyses demonstrate that effective transfer depends on how skills are represented, favoring concise, operation-oriented supervision over generic image descriptions or explicit reasoning traces. These results support visual skill internalization as a complementary direction for medical multimodal learning, in which clinically structured visual expertise is transformed from external knowledge into reusable model capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.