acceptodds
Under review as a conference paper at ICLR 2027

Can Gray-Box Rival White-Box? Anticipatory Consistency Propagation with Top- Candidates for LLM Hallucination Detection

Abstract

Hallucinations in large language models are a major obstacle to their applications. While white-box detectors achieve strong performance by exploiting internal model states, such information is inaccessible for closed-source LLMs. Gray-box methods, on the other hand, offer a practical alternative by leveraging only API-accessible Top- candidates and their probabilities, but their performance remains limited. In this paper, we seek to bridge this gap by introducing rich external semantic representations into gray-box detectors. However, since external embedding models and the target LLM are trained independently, an inherent information gap could exist between the external semantic space and the target model’s generation dynamics. To address this, we propose **Prop-K**, a consistency-guided propagation method for gray-box hallucination detection. Prop-K captures the predictive consistency of the target model’s generation process at two complementary levels: **local agreement consistency** measures how Top- candidates support or conflict with the current Top-1 prediction at each step, while **global anticipatory consistency** captures whether early predictive expectations align with or deviate from what the model generates later. Together, these two consistency signals guide semantic propagation throughout the generated sequence, potentially aligning external semantic embeddings with the target model’s predictive behavior. Empirically, Prop-K consistently outperforms existing gray-box methods with an average gain in AUROC of 0.04–0.07 **across 7 target LLMs and 4 benchmarks**. Notably, despite having no access to internal model states, Prop-K matches and, in some cases, even surpasses the strongest white-box detector with such privileged access.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.