MedCode: Medical Visual Reasoning beyond Predefined Toolkits
Abstract
Medical vision language models (VLM) have benefited from explicit reasoning, but most existing approaches reason over a visual representation that is encoded only once. Recent tool augmented methods allow models to revisit the image, yet their analyses remain limited to procedures defined in advance. In this paper, we present MedCode, a framework that enables medical VLMs to construct question specific image analyses through executable code. Instead of selecting from a predefined toolkit, MedCode writes programs using general image processing and numerical operations. We further develop MedCodeLab, a shared general purpose environment that executes these programs and returns textual results and derived images for subsequent reasoning. To teach this capability, we construct MedCode-50k, which contains 49,817 verified multi turn trajectories across eight imaging modalities and four question categories. We train an 8B model through trajectory based supervised fine tuning followed by agentic reinforcement learning on questions selected according to rollout difficulty. On the in distribution test set, MedCode raises the accuracy of its backbone from 75.2% to 82.8%. It also generalizes to four unseen data sources, where it improves average accuracy from 76.2% to 85.5%. MedCode consistently outperforms crop based reasoning and predefined image operations, demonstrating the value of constructing flexible analyses for medical visual reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.