PharmEval: A Diagnostic Benchmark for Pharmacological Knowledge and Reasoning of Large Language Models
Abstract
Existing pharmaceutical LLM benchmarks primarily focus on accuracy, making it difficult to determine whether model failures arise from insufficient knowledge, ineffective information use, or unstable reasoning. We introduce PharmEval, a pharmaceutical benchmark comprising 6,174 questions across 9 task families and 47 subcategories. Beyond conventional accuracy-based evaluation, PharmEval assesses explanation quality and incorporates four diagnostics of pharmaceutical knowledge use, including web-based knowledge retrieval, compositional knowledge use, counterfactual sensitivity, and response modality variation. These studies examine how external information and explicit premises support knowledge use, and whether models maintain or correctly update their judgments as conditions and response formats change. We evaluate a diverse set of LLMs and conduct an in-depth analysis of representative models. Our results show that benchmark accuracy alone does not fully reflect how reliably models use pharmaceutical knowledge. High accuracy can coexist with weak open-ended performance and failures to update judgments correctly when key conditions change. These findings suggest that a key challenge for current LLMs lies not only in acquiring relevant knowledge, but also in applying it consistently and reliably in context. By evaluating models from both performance and behavioral perspectives, PharmEval provides a more nuanced characterization of their capabilities and limitations in pharmaceutical reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.