FALD: Fine-grained Adaptive Latent Detection for LLM Hallucination
Abstract
Hallucination detection is essential for the reliable deployment of large language models, yet latent-space methods often treat all tokens uniformly during intervention and fix the readout at the final layer, which can dilute token-localized hallucination evidence and limit feature separability. We propose FALD, a fine-grained adaptive latent detection framework that addresses both limitations. FALD first learns a token-level probe—trained under multiple-instance weak supervision from sequence-level labels only—to identify hallucination-prone tokens and modulate latent intervention per token, so that the separator focuses on the most informative evidence. It further performs layer-wise probing to select the most hallucination-sensitive readout layer rather than assuming the final layer is optimal. By combining token-aware intervention with adaptive feature-layer readout while keeping the backbone frozen, FALD learns more discriminative latent representations. Across four QA benchmarks and three LLM families (LLaMA, Qwen, Mistral), FALD consistently outperforms the strongest latent baseline (TSV) by 2.7-7.6% AUROC, and surpasses multi-sample methods by >24% AUROC at a fraction of the inference cost, highlighting the value of fine-grained training signals and layer-adaptive readout.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.