acceptodds
Under review as a conference paper at ICLR 2027

HD-Fuse: Single-Pass Hallucination Detection via Confidence Reshaping with Conformal Likelihood-Ratio Fusion under Logprob-Only APIs

Abstract

Hallucination detection under API logprob-only contracts is practically important but constrained: production LLM APIs expose only final-token logprob distributions, while hidden states, attention weights, and repeated sampling are unavailable or expensive. We present a systematic empirical study of single-pass, logprob-only hallucination detection for short-form QA ( tokens) and identify a validated recipe that outperforms individual signals: extract a 5-dimensional confidence vector, model the test as a Gaussian likelihood ratio (LDA/QDA), fuse complementary likelihood-ratio members (including thermodynamic trajectory statistics of the top- distribution) via a calibrated cross-validated rank ensemble, and attach a split-conformal -value with a finite-sample false-positive-rate guarantee, at zero extra passes. On real public benchmarks the recipe attains the best cross-cell mean AUC at one pass versus five; we confirm the signal and false-positive control on a production API (DeepSeek), so the logprob-only claim does not rest on local quantised weights. Ablation attributes the gain to signal decorrelation. Scope and prerequisites (short answers; a per-(model, benchmark) calibration fold) are stated up front, so the contribution is a validated deployment recipe plus an interpretable diagnostic framework; longer-form generation is future work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.