PedMed-Bench: A Multi-turn Benchmark for Large Language Models on Pediatric Medication Effectiveness and Safety
Abstract
Evaluating Medical Large Language Models (LLMs) in complex clinical scenarios has gained increasing attention. Yet most existing benchmarks emphasize general or adult-centered medicine, making them poorly aligned with pediatric care, where age-dependent physiology, longitudinal disease evolution, and medication safety require specialized reasoning. Specifically, current evaluation systems regarding pediatric care are still facing severe data scarcity challenges, thereby hindering the evaluation and development of pediatric Medical LLMs. Motivated by this, we introduce a novel benchmark (PedMed-Bench) to address the above challenges. Further, we conduct a systematic evaluation involving numerous representative LLMs on PedMed-Bench. The results suggest that contemporary LLMs still struggle to accurately predict adverse drug events (ADEs) and suspected drugs. Even though models equipped with reasoning capabilities demonstrate notable enhancements, their absolute performance remains constrained, with the highest ADE extraction F1 score reaching only 12.89%. To the best of our knowledge, PedMed-Bench is the first fine-grained, event-driven benchmark for real-world pediatric medication evaluation. The benchmark is released in both Chinese and English, with human-verified translation validated for cross-lingual equivalence. We hope it can serve as a helpful resource, contributing to the ongoing efforts in developing more reliable clinical AI tools for pediatric healthcare.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.