Narrative Fidelity: Beyond Factuality in Evaluating Clinical Dialogue Summarization
Abstract
Clinical notes are usually evaluated as summaries, with ROUGE, BERTScore, or other overlap with a reference note. These metrics reward factual completeness but say little about whether a note preserves the patient's own framing, concerns, uncertainty, and perspective. We call this property narrative fidelity and introduce a six-dimension annotation taxonomy for it. We apply the taxonomy to notes from six systems (two current Claude models and four published baselines) on the ACI-Bench clinical dialogue corpus, and report two findings. First, narrative fidelity agrees with ROUGE on system ranking (Spearman ) but not on individual cases. Across cases and systems, its correlation with ROUGE-1 is moderate (– across core dimensions), well below the correlation among the ROUGE variants themselves (–), so the taxonomy captures variance that ROUGE does not. ROUGE tracks Clinical/Factual Content most closely () and Uncertainty/Epistemic Stance least (). BERTScore correlates less still (– per core dimension) and barely separates the systems: mean BERTScore varies by across systems whose narrative fidelity ranges from to . Second, we test sensitivity to a naturally occurring degradation, comparing ASR and human transcripts of the same 28 encounters. Aggregate narrative fidelity and ROUGE-1 change little, but Patient Concerns & Priorities declines (, paired -test; , Wilcoxon), and the decline is spread across cases rather than driven by outliers. This result does not survive correction for the six dimensions tested, so we treat it as a hypothesis for replication. It suggests that transcript quality can erode specific narrative content when overall note quality looks unchanged. Inter-annotator agreement is strong on three dimensions and lower for Perspective/Attribution, and we discuss what these results imply for evaluating AI scribes beyond factuality checks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.