acceptodds
Under review as a conference paper at ICLR 2027

Clinical Hallucination Detection with Coverage-Gated Ontology Checks

Abstract

Large language models (LLMs) that write clinical text can state false facts in a fluent, convincing form. A common safeguard is a second LLM, an auditor, that checks each claim by calling external knowledge tools. In practice this is fragile: auditors misuse the tools or misread their replies, and some models, including the medical model we test, cannot call tools at all. We propose coverage-gated ontol- ogy checking, which takes the knowledge lookup out of the auditor’s hands. Code first decides whether a clinical ontology can speak to a claim. If it can, the ontol- ogy checks the claim, and the auditor reads the results and gives a single verdict, without ever calling a tool. We evaluate the approach on fourteen auditors from seven model families, from 3B to 128B parameters. On a clinical hallucination benchmark it improves every auditor, raising the Matthews correlation coefficient by 0.09 to 0.30 over the same auditor without checks. The gains are largest for weaker models, and with checks a 7B auditor matches 70B and 120B auditors without them. The approach also works for models that cannot call tools, and it outperforms tool calling for most models that can. Because coverage is decided before the auditor runs, it shows in advance where the ontology can help.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.