MedRoAM: Improving Rollout Robustness and Intervention Sensitivity in Medical Language World Models
Abstract
Medical language world models (LWMs) use large language model (LLM) backbones to predict how a patient's state evolves given the clinical history and an intervention. Despite their promise, medical LWMs remain at an early stage of development, and their limitations are not yet well understood. Through a series of preliminary analyses, we identify two key shortcomings: (1) limited rollout robustness, reflected in a substantial performance drop from teacher-forced to autoregressive simulation, and (2) weak intervention sensitivity, where predictions barely change even when the probing intervention is removed. To address these shortcomings, we propose MedRoAM, which jointly optimizes two simple-yet-effective training objectives: Privileged Self-Distillation (PSD) and Intervention Response Alignment (IRA). PSD uses a frozen teacher conditioned on ground-truth histories to guide a student trained on a mixture of clean and noisy rollout histories, encouraging robustness to accumulated errors while anchoring predictions to the initial model. IRA incorporates pharmacological knowledge to align numerical response estimates across the original intervention, its removal, and a similarly acting substitute. Extensive experiments across several LLM backbones show that MedRoAM improves autoregressive rollout accuracy on in-distribution and out-of-distribution data while largely preserving teacher-forced performance. Evaluations on 14 supervised and 7 held-out intervention-response probing pairs further demonstrate improved intervention sensitivity and pharmacological consistency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.