CodeLancet: Internalizing Physician Skills for Diagnosis via Agentic Coding
Abstract
Vision–language models (VLMs) now approach specialists at recognizing what a medical image shows. Yet many diagnoses must be examined rather than recognized. On these, their correctness is often hollow. On ejection fraction, which must be measured by tracing the left ventricle through the cardiac cycle, the most accurate frontier VLMs miss 74–89% of failing hearts and score well only because most hearts are normal. What they lack is less medical knowledge than a means of acting on the image, so they fall back on the population prior. Can a VLM learn to examine the way a physician does? Our key insight is that executable code turns examination into a sequence of executed and inspectable acts on the image. Unlike a fixed toolset, code lets the model check and correct its own evidence, compose operations no designer anticipated, adapt its effort to the case and make perception itself an act, choosing where and how finely to look. We present CodeLancet, a 4B agent that examines medical images through executable code. Skill-guided, rubric-audited demonstrations and outcome-based GRPO first teach it to conduct the examination. Yet a trajectory-level outcome reward credits every step of a correct examination alike, and over a third of these contain critical wrong steps. Focal-PPO instead credits each decision by the change in expected outcome it causes, estimated by branching rollouts around the decisions where failed examinations first went wrong, and the second stage cuts this share to a tenth. Across six benchmarks in fundus photography, echocardiography and gigapixel pathology, CodeLancet outperforms six frontier VLMs on every headline metric: at comparable overall accuracy it detects 72.5% of failing hearts, it nearly halves referral errors in retinopathy, and on whole-slide melanoma images it issues about a quarter fewer false relapse alerts, with a margin that widens as the answer depends more on examination than on recognition. Blinded specialists endorse 80–86% of its steps. Our results suggest that diagnosis is learned not only by seeing better but also by acting on what is seen, and that code can be to a model what the lancet is to a physician.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.