acceptodds
Under review as a conference paper at ICLR 2027

RadThinking: Teaching Vision-Language Models to Think the Way Radiologists Are Taught

Abstract

Radiology education teaches a way of thinking, not a set of answers. A trainee learns to screen each region of an organ, characterize each finding, compare it with earlier scans, and apply a clinical guideline, a rule over the finding's representation: its size, enhancement, change, and a few more. Vision-language models are trained on images paired with answers and judged by the answer alone; the representation in between goes unchecked. We present RadThinking, which turns this way of thinking into supervision: the representation a radiologist is taught to build becomes the representation the model learns to write. Each step of reading a scan becomes a question with a reference answer, grounded in a 3D bounding box; questions link into chains; five clinical guidelines become decision tables, explicit rules from a finding's representation to its category. The representation is grounded, so each answer in it is scored at its box on the scan, and explicit, so a radiologist can check it as they check a trainee. RadThinking holds 122,127 questions over 5,365 abdominal CT scans of 2,679 patients from two countries. Across eight models, the best gets every answer right in only 23% of chains, numerical errors aside, and turning thinking on brings no consistent gain. Looking inside open-weight models shows why. Once its intermediate answers are written, Qwen3.8-27B follows that text, not the images, right or wrong. Before it writes, a linear probe on its hidden state predicts the final answer about 9 points more often than the model does. The model holds more than it writes and follows what it writes. So the representation it writes must be correct, and RadThinking supervises every answer in it. Trained on it, a 4B model outperforms zero-shot GPT-6, Claude Opus 5, and Gemini 3.8 Flash: 64% versus 52% on individual questions and 30% versus 23% on chains. Radiology education already wrote this representation down. A small model that learns it outperforms much larger models that have not.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.