acceptodds
Under review as a conference paper at ICLR 2027

CodeLancet: Internalizing Physician Skills for Diagnosis via Agentic Coding

Abstract

Vision–language models (VLMs) now approach specialists at recognizing what a medical image shows. Yet many diagnoses must be examined rather than recognized. On these, their correctness is often hollow. On ejection fraction, which must be measured by tracing the left ventricle through the cardiac cycle, the most accurate frontier VLMs miss 74–89% of failing hearts and score well only because most hearts are normal. What they lack is less medical knowledge than a means of acting on the image, so they fall back on the population prior. Can a VLM learn to examine the way a physician does? Our key insight is that executable code turns examination into a sequence of executed and inspectable acts on the image. Unlike a fixed toolset, code lets the model check and correct its own evidence, compose operations no designer anticipated, adapt its effort to the case and make perception itself an act, choosing where and how finely to look. We present CodeLancet, a 4B agent that examines medical images through executable code. Skill-guided, rubric-audited demonstrations and outcome-based GRPO first teach it to conduct the examination. Yet a trajectory-level outcome reward credits every step of a correct examination alike, and over a third of these contain critical wrong steps. Focal-PPO instead credits each decision by the change in expected outcome it causes, estimated by branching rollouts around the decisions where failed examinations first went wrong, and the second stage cuts this share to a tenth. Across six benchmarks in fundus photography, echocardiography and gigapixel pathology, CodeLancet outperforms six frontier VLMs on every headline metric: at comparable overall accuracy it detects 72.5% of failing hearts, it nearly halves referral errors in retinopathy, and on whole-slide melanoma images it issues about a quarter fewer false relapse alerts, with a margin that widens as the answer depends more on examination than on recognition. Blinded specialists endorse 80–86% of its steps. Our results suggest that diagnosis is learned not only by seeing better but also by acting on what is seen, and that code can be to a model what the lancet is to a physician.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.