MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning
Abstract
Medical reasoning can reach a plausible answer while overlooking an unsafe or unsupported intermediate step. Process-Level Reward Models (PRMs) can verify such steps, yet existing PRM benchmarks largely target mathematics rather than patient-specific clinical errors. We introduce MedPRMBench, a step-level medical reasoning benchmark with 14 clinical error types, including five new and two clinically re-scoped types, and rule-defined severity grades. Built from seven medical QA sources via structured error construction and review by 86 physicians from Grade-A tertiary hospitals, it provides 6,500 evaluation questions, 13,000 reasoning chains, and 113,910 step labels, plus a leakage-clean 6,870-question training split. The best of 20 prompted critics reaches only 75.4% PRMScore, and critics differ sharply in error-type profiles and operating points. A Qwen3-8B PRM trained on the split reaches 86.6% PRMScore and 89.0% strict first-error localization; in matched ablations, physician review raises error recall by 6.55 points and cuts clean-chain false alarms from 5.6% to 2.1%. The training signal replicates across model families; without further training, the PRM transfers to ProcessBench and PRMBench, exceeding six of nine math-trained PRMs on each.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.