Palimpsest: Bidirectional Scoring for Span-Level Hallucination Detection
Abstract
Large language models can produce "hallucinated" statements that sound fluent but are ungrounded in reality, motivating reliable detection. Even though support for a span of statement can depend on words that follow it, many evaluators score text left to right. Motivated by this shortcoming, we propose Palimpsest, a span-level detector that uses masked diffusion language models (MDLMs), which evaluates the factuality of a span of text using its context from both sides. We compare Palimpsest to autoregressive (AR) models with AUROC and find a) on BUMP with human-annotated spans, 8B MDLM scorers match the strongest tested AR model, which has 3.9 times as many nominal parameters, and b) on LLM-AggreFact they lag behind slightly while being more parameter-efficient. On HalluEntity, LLaDA-8B-Base outperforms the strongest tested AR scorer, a 31B checkpoint. Thus, Palimpsest can rival larger AR-based methods for span-level hallucination detection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.