LUCiD: Mitigating Hallucinations in Large Language Models with Learned Unilayer Contrastive Decoding
Abstract
Large language models remain prone to hallucination, generating fluent but factually wrong or unsupported content. This is often traced to models falling back on pretraining priors instead of grounding each token in the available context, a failure that is costly in reasoning and knowledge-intensive tasks. Contrastive decoding methods such as DoLa improve reasoning and factuality by contrasting a model’s final-layer distribution against a premature earlier-layer distribution. They are training-free, so the premature signal is taken directly from an untrained early exit and the contrast strength is fixed across tokens, which is brittle and often lowers accuracy on tasks where earlier layers are already well calibrated. We introduce LUCiD (Learned Unilayer Contrastive Decoding), a learned contrastive decoding method. LUCiD attaches one lightweight probe to a single middle layer and reuses the model’s own frozen final RMSNorm and LM head to extract a premature distribution. This distribution is contrasted against the final layer, scaled by a per-token gate predicted from the last hidden state, under an adaptive plausibility constraint. Across five models from 0.6B to 12B and ten benchmarks, under both greedy decoding and sampling, LUCiD outperforms vanilla decoding, DoLa and SLED, in most scenarios.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.