acceptodds
Under review as a conference paper at ICLR 2027

HalluTracer: Pre-Decoding Truthfulness Prediction via Depth-Averaged Probe-Logit

Abstract

Internal-state probes enable truthfulness prediction before a large language model generates an answer. When detectors change both the layers they read and the rules used to combine them, the source of improved prediction becomes difficult to identify. We separate these choices and find that retaining more layers improves prediction even under fixed equal weighting. An exact Fisher-ratio decomposition explains why the additional benefit of linear reweighting is limited on these probe scores: information in layer-wise differences largely overlaps with that captured by the depth mean. Estimating additional weights can then offset this small benefit when the data used to fit them are limited. These findings motivate our proposed method HalluTracer, which averages layer-wise probe logits to predict truthfulness before decoding. Across six models and four benchmarks, HalluTracer achieves the highest AUROC in 23 of 24 model–benchmark pairs among the compared methods. The results support using evidence from across the network without requiring a correspondingly more flexible aggregation rule, clarifying the distinct roles of layer selection and weighting in pre-decoding detection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.