acceptodds
Under review as a conference paper at ICLR 2027

MedEvo: Medical Reasoning Evolution through Tag- Guided Navigation

Abstract

Medical multimodal reasoning has advanced through teacher-generated rationale distillation, verifier- or reward-guided learning, and on-policy supervision over student-generated rollouts. Although these paradigms can exploit mistakes through rejection, reward, preference, process, or token-level signals, explicitly recovering valid reasoning states from failed medical trajectories and converting localized corrections into reusable, student-relative supervision remains underexplored. Moreover, general-domain correction mechanisms do not directly address medical failures grounded in modality-specific visual evidence, anatomical context, evidence integration, and differential diagnosis. We propose MedEvo, a failure-conditioned framework for evolving medical multimodal reasoning through localized trajectory repair and feedback-driven data evolution. At the trajectory level, MedEvo uses step-level uncertainty to propose a candidate failure region, employs a gold-blind auditor to identify a safe rollback boundary and generate a taxonomy-guided minimal recovery hint, and lets the original student preserve the valid prefix while regenerating only the erroneous suffix. Repair-aware learnability filtering then retains answer-correct, locally learnable trajectories that remain compatible with the student policy. At the data level, a fine-grained medical taxonomy characterizes checkpoint-specific capability gaps and guides the selection of targeted supervision for the next evolution round. Using supervised fine-tuning alone, MedEvo improves Qwen3.5-9B on MedXpertQA-MM from 47.40% to 54.50% and raises its seven-benchmark mean from 65.57% to 69.12%, a gain of 3.55 percentage points. Experiments on 9B, 35B, and 397B models show gains in mean benchmark accuracy across scales, while round-wise comparisons examine feedback-driven data evolution. These results establish localized failure repair and structured feedback as a practical step toward recursive self-improvement in medical reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.