MedAgentWorld: Privileged On-Policy Self-Distillation for World-Model-Augmented Clinical Agents
Abstract
Medical large language models perform well on static case questions but are rarely evaluated on the sequential decisions of acute care. We present MedAgentWorld (MedWA), a world-model-augmented reinforcement learning framework for multi-turn clinical planning. We curate MedWA-118K and MedWA-Bench as time-ordered patient trajectories and build MedWA-Gym, where an agent acts on currently available information and receives replayed observations. Because historical records lack responses to plausible actions that were not performed, MedWA-WM predicts missing next observations. To make the world model robust to histories containing its own predictions, we introduce privileged on-policy self-distillation (OPSD): a student rolls out on its generated histories, while a teacher uses training-only access to the target observation, subsequent examination results, and final diagnosis to provide distillation targets. At inference, the student conditions only on the visible patient history and selected action. We train MedWA-Agent with supervised fine-tuning and reinforcement learning on replayed and world-model-generated branches. On MedWA-Bench, the agent reaches 44.92% diagnosis accuracy, compared with 42.73% without the world model, and MedWA-WM improves held-out next-observation prediction over prompted language-model baselines. These results support world-model-assisted learning for interactive clinical planning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.