REVEAL: Evaluating Medical LLMs with Risk-Grounded Virtual Patients
Abstract
Medical LLMs are commonly evaluated with single-turn clinical questions, yet real online consultations often unfold through missing histories, delayed disclosure, corrections, and evolving risk information. We present **REVEAL**, a risk-aware virtual-patient framework for multi-turn medical LLM safety evaluation on **753 real text-only online consultations**. REVEAL builds source-extracted patient profiles and compares three patient simulators: a **prompt-only patient simulator** that conditions a patient-role prompt on the extracted profile, a **risk-hiding virtual patient** that withholds latent risk attributes unless elicited, and a **DialoguePlan-based risk-replay virtual patient** that reproduces source-derived risk-emergence trajectories. We evaluate patient fidelity before downstream safety: in a 753-case human arena with three annotators, risk-replay is selected as the closest source-style transcript in 71.2% of cases (Fleiss’ κ = 0.62). We then test five evaluated doctor models under both answer-first and history-taking settings with a rubric-based medical-risk judge. Under a source-matched but non-equivalent single-turn protocol, all three multi-turn simulators produce higher severe-failure rates than the single-turn setting (32.14–38.18% versus 22.67%), indicating complementary interaction-level safety failures. Mechanism checks show that risk-hiding and risk-replay induce distinct disclosure regimes: 3.60% versus 99.49% tracked-risk disclosure. An exploratory challenge-behavior perturbation does not increase unsafe outcomes overall, but information contradiction exposes selected boundary failures. REVEAL therefore separates three quantities that are often conflated in medical LLM evaluation: patient realism, risk-information emergence, and downstream doctor safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.