MedEvo: Medical Reasoning Evolution through Tag- Guided Navigation
Abstract
Medical multimodal reasoning has advanced through teacher-generated rationale distillation, verifier- or reward-guided learning, and on-policy supervision over student-generated rollouts. Although these paradigms can exploit mistakes through rejection, reward, preference, process, or token-level signals, explicitly recovering valid reasoning states from failed medical trajectories and converting localized corrections into reusable, student-relative supervision remains underexplored. Moreover, general-domain correction mechanisms do not directly address medical failures grounded in modality-specific visual evidence, anatomical context, evidence integration, and differential diagnosis. We propose MedEvo, a failure-conditioned framework for evolving medical multimodal reasoning through localized trajectory repair and feedback-driven data evolution. At the trajectory level, MedEvo uses step-level uncertainty to propose a candidate failure region, employs a gold-blind auditor to identify a safe rollback boundary and generate a taxonomy-guided minimal recovery hint, and lets the original student preserve the valid prefix while regenerating only the erroneous suffix. Repair-aware learnability filtering then retains answer-correct, locally learnable trajectories that remain compatible with the student policy. At the data level, a fine-grained medical taxonomy characterizes checkpoint-specific capability gaps and guides the selection of targeted supervision for the next evolution round. Using supervised fine-tuning alone, MedEvo improves Qwen3.5-9B on MedXpertQA-MM from 47.40% to 54.50% and raises its seven-benchmark mean from 65.57% to 69.12%, a gain of 3.55 percentage points. Experiments on 9B, 35B, and 397B models show gains in mean benchmark accuracy across scales, while round-wise comparisons examine feedback-driven data evolution. These results establish localized failure repair and structured feedback as a practical step toward recursive self-improvement in medical reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.