acceptodds
Under review as a conference paper at ICLR 2027

When Long Contexts Break Faithfulness: Evaluating Medical Hallucination under Adversarial Evidence

Abstract

Real-world medical dialogues require LLMs to reason over multiple sources, from retrieved documents and conversation history to family records, test reports, and personal health data. Existing medical hallucination benchmarks mainly target factual errors, but overlook faithfulness errors: misattributing external, historical, or non-user information to the current user, potentially leading to misleading user-specific medical advice. We introduce MLoHa, a long-context medical faithfulness benchmark with 2,700 items covering three objective tasks (choice, judgment, and best_answer), a fill-in task, and a five-dimensional Rubric QA task. Stratified into four length tiers (LA–LD), MLoHa evaluates 14 large language models under realistic multi-source medical contexts. Across all task formats, faithfulness degrades monotonically with context length: choice accuracy drops by 22.08 percentage points from LA to LD, and the five-task average drops by 13.22 points. An objective-only multi-factor ANOVA further shows that context length is the strongest explanatory factor (partial , ), outweighing medical scenarios and hallucination triggers. The only significant interaction is length medical scenario (), indicating that long contexts reshape how different medical risks surface, while ranking differences across tasks show that user-grounded faithfulness is not captured by a single evaluation format. We hope MLoHa can provide a reusable benchmark and analysis framework for the community to evaluate user-grounded medical faithfulness in long-context, multi-source settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.