acceptodds
Under review as a conference paper at ICLR 2027

CoEvoMed: Co-Evolving Supervision Principles and Medical Vision–Language Models

Abstract

Training medical vision–language models requires supervision that adapts as the learner's capabilities change across tasks. Existing pipelines typically rely on repeated manual error analysis and task-specific training data construction, making this adaptation costly. We introduce CoEvoMed, a co-evolutionary framework that couples a medical vision–language learner with a multi-agent supervision process integrating the evolution of construction principles with curriculum allocation. These reusable principles define question intent, guide how responses draw on medical evidence, and clarify conceptual distinctions in training examples. A revision agent updates these principles in response to the learner's evolving task-specific capabilities, guided by comparisons of factual consistency and task alignment between examples generated from the same medical evidence. Under fixed budgets for each task, curriculum allocation adjusts construction effort according to the learner's capability profile and learning progress. Using CoEvoMed, we train an 8B model supporting both 2D and 3D medical tasks. It surpasses Hulu-Med-7B by 1.15–4.80 percentage points on four 2D VQA benchmarks and achieves higher RaTEScore on MIMIC-CXR and IU-Xray, while delivering strong performance on 3D recognition, question answering, and report generation. Results at 32B further support the scalability of the framework.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.