acceptodds
Under review as a conference paper at ICLR 2027

Fully Open Meditron: Auditable and Reproducible Development of Clinical LLMs

Abstract

Clinical decision-support LLMs are either proprietary or open-weight, without the training data and development pipelines required for end-to-end reproducibility and auditing. We investigate how much of the gap between fully-open and open-weight clinical specialists is closeable under full-openness constraints. We introduce MeditronFO (Meditron Fully Open), to our knowledge, the first fully-open pipeline for clinical LLM finetuning. The pipeline creates three new synthetic datasets (340k QA pairs) from clinician-authored adversarial vignettes, clinical practice guidelines and public medical QA, using open-weight gpt-oss-120b for rejection-sampling distillation with pass@8, with generation prompts co-authored by a panel of four clinicians. We apply the pipeline to five fully-open base models from three families and one open-weight control (Gemma-3-27B). On the average of five medical benchmarks, Apertus-70B-MeditronFO reaches 54.56%, against 48.35% for Apertus-70B-Instruct and 63.75% for MedGemma-27B, closing 40% of the gap; the Gemma-3-27B-MeditronFO control reaches 58.65%. We also introduce AutoMOOVE, an automated open-ended clinician preference benchmark, whose judgments agree with blinded clinicians' preferences (). The Gemma-3-27B-MeditronFO control is level with or ahead of MedGemma-27B on this benchmark (57.1% win rate under a three-judge majority vote, above parity under two of three judges). We release the data, code and model weights as public artifacts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.