acceptodds
Under review as a conference paper at ICLR 2027

Risk-Controlled Detection of Hallucinated References with Aggregated Detectors

Abstract

Large language models frequently fabricate academic references that appear authentic but correspond to no real publication. Existing detectors typically rely on a single aspect of verification and offer no formal guarantees on their error rates. We design a risk-controlled framework that assigns each reference a continuous suspicion score by aggregating heterogeneous verification signals, and then labels it as valid, hallucinated, or deferred to a human, controlling the false positive and false negative rates with as few deferrals as possible. Our guiding principle is that validity comes from calibration and efficiency from aggregation. We calibrate two thresholds with per-class conformal quantiles of the suspicion score, which delivers finite-sample control at user-specified levels. We then show that the deferral rate depends on the score only through its ROC curve, and that the posterior over detector outputs is the optimal aggregator. Experiments on the HallMark benchmark and on StatRef, a new dataset of statistics references, confirm that calibration secures validity and aggregation secures efficiency, with recalibration restoring control under domain shift.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.