acceptodds
Under review as a conference paper at ICLR 2027

AWARD: Answer-Level Weighted Aggregation of Retrieval-Free Diagnostics for Hallucination Detection

Abstract

Hallucination in large language model (LLM) outputs is often flagged by inference-time signals that are then used directly as answer-level scores, despite measuring a different quantity. This paper presents AWARD, a retrieval-free hallucination detector that separates a formal sequential source diagnostic from a supervised answer-level score. The core signal is Speculative Cross-Model Divergence (SCMD), the pointwise log ratio between an instruction-tuned model and its own pretrained base. Under the base-sampling null, the SCMD product is an e-process satisfying Ville's anytime type-I bound. SCMD is one of four inference-time primitives, joined by primary-model entropy, Attention-Grounding Entropy (AGE), and Semantic Refutation via NLI (SRN). Across five model families and three benchmarks (15 cells), nested cross-validation yields mean AUROCs of , , and on HaluEvalQA, MatSciQA, and TruthfulQA, respectively. In comparisons of 10-seed paired-bootstrap with the strongest of DoLa, SelfCheckGPT, and Semantic Entropy, AWARD achieves mean , with positive differences in cells; two survive Holm–Bonferroni correction in family . AWARD wins cells with mean against the stronger of two classifier- and label-matched supervised fusions. Split-conformal calibration provides marginal false-positive-rate control under exchangeability. At the lowest non-degenerate tier, , cells meet the seed-averaged target and meet an item-clustered upper-confidence-bound criterion; the tier is degenerate in every cell because the calibration sets are too small. All four primary feature blocks have positive mean contributions in leave-one-block-out ablations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.