Evidence-to-Capability Matching for Medical Training Data Synthesis
Abstract
Synthetic training data for medical large language models (LLMs) draws on heterogeneous sources, but topical relevance does not guarantee that a source supports the capability a target task requires. We introduce Evidence-to-Capability Matching (\method), a framework that profiles target capabilities, identifies source-supported evidence, and specifies requirements for constructing training questions. Its refinement component, Medical Profile-Routed Agentic Refinement (\refiner), coordinates source-constrained revision and verification for medical validity and capability alignment. On Qwen3-8B, task-focused corpora improve their primary benchmarks for evidence-intensive reasoning, clinical decisions, and knowledge coverage by 2.69, 6.68, and 1.46 percentage points, respectively. Ablations of six guidance families show that full guidance performs best in both tested settings, while the most consequential guidance differs by task: knowledge targeting for clinical decisions and evidence structuring for evidence-intensive reasoning. Clinical-decision-oriented training also improves on a benchmark excluded from synthesis and training, and a cross-task mixture improves over the base model on all four benchmarks. These results suggest that effective medical supervision depends on matching what a task requires to what its sources can support.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.