acceptodds
Under review as a conference paper at ICLR 2027

What Do Hallucination Benchmarks Measure? Exploratory Factor Analysis for Detection and Benchmark Diagnosis

Abstract

Hallucination detectors are usually evaluated in the same way, although they capture different properties of generated text. We present an unsupervised measurement framework that treats hallucination severity as a latent trait and uses a battery of indicators to distinguish input-conflicting (IC) hallucinations, which are unsupported by a source document, from context-conflicting (CC) hallucinations, inconsistent with the generated text itself. Within each IC and CC battery, heterogeneous indicators are rank-normalised and combined using a bifactor exploratory factor model, yielding a continuous score without hallucination labels, human-written reference outputs, or API calls. We evaluate seven benchmarks covering summarisation, source-grounded question answering, and document-level contradiction. On five IC-only benchmarks, source-grounding scores discriminate hallucination labels better than output-only coherence scores. Detection performance is competitive but varies markedly across datasets. Furthermore, Factor loadings show how strongly semantic similarity and source consistency contribute to the shared score. Across the benchmarks, detection is easier when faithful and hallucinated outputs differ in how much source wording they reuse, although our score provides information beyond this difference. After stratifying examples with similar levels of lexical novelty, our method retains more discriminative signal on benchmarks constructed from LLM-generated hallucinations than on extractive summarisation benchmarks, while generally adding predictive information beyond novelty alone. These findings suggest that current leaderboards can conflate factuality sensitivity with extractiveness. We therefore propose lexical-stratified evaluation and incremental-validity testing as complementary checks for hallucination benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.