READ: Combining Reader Evidence for Non-Intrusive Black-Box Hallucination Detection
Abstract
Hallucination detection supports reliable use of large language models (LLMs), but many existing methods require internal access or repeated interaction with the target model. We study non-intrusive black-box hallucination detection (NI-BHD), which assesses previously generated responses without either requirement. An accessible model, termed a reader, can provide detection signals through teacher forcing on the question and observed answer tokens. Our analysis reveals a reader suitability mismatch: model size, answering ability, and reader–target family match do not reliably guide reader selection. Reader rankings also vary across targets and benchmarks; readers provide complementary evidence. Accordingly, we derive an information-theoretic characterization of a reader’s contribution beyond the selected evidence via optimal population log-loss reduction. We therefore propose reader evidence aggregation for detection (READ), a supervised framework selecting and combining multiple readers. READ selects readers by cross-validated log-loss reduction and integrates their hidden-state probe scores and token-likelihood features using a fixed additive logistic model. Experiments on 6 target LLMs and 7 benchmarks achieve 0.911 cell-macro AUROC, improving over the strongest published baseline under the same access constraints by 0.116. Despite strictly less access, READ also exceeds the per-environment best of 7 published white-box detectors in 19 of 28 matched environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.