DCLED: Dynamic Confidence-Aware Layer-wise Evolution Decoding for Factuality-Oriented Candidate Scoring
Abstract
Large language models often assign high likelihood to fluent but factually incorrect candidate answers in knowledge-intensive benchmarks. We introduce DCLED (Dynamic Confidence-Aware Layer-wise Evolution Decoding), a training-free inference-time method for factuality-oriented multiple-choice and ranking tasks. DCLED combines three mechanisms inside iterative logit evolution: an early-layer consensus anchor with a mature–latent contrastive term, a confidence gate that applies evolution only at uncertain token positions, and a confidence boost that scales the evolution step with top-token concentration. Across TruthfulQA and HotpotQA with five Llama, Qwen, and Mistral checkpoints, DCLED obtains the highest TruthfulQA MC2 on all five checkpoints and the highest HotpotQA ranking accuracy on four of five. The largest gain occurs with Llama-3.2-1B on HotpotQA, where accuracy increases from 52.99% to 78.27% over vanilla decoding (47.7% relative) and from 57.41% to 78.27% over SLED. HotpotQA gains over SLED are statistically significant for four of five checkpoints, whereas TruthfulQA MC2 gains over SLED (+0.013 to +0.055) are not significant after multiple-comparison correction. DCLED adds 12–64% latency relative to SLED and 15–185% relative to vanilla decoding. Ablations identify confidence gating as the dominant mechanism on HotpotQA, while the contrastive term and confidence boosting provide checkpoint-dependent refinements. DCLED is therefore best viewed as a parameter-preserving method for confidence-conditioned candidate scoring when hidden states are accessible, not as a universal replacement for factuality-oriented decoding baselines. Code is available at https://anonymous.4open.science/r/DCLED-Dynamic-Confidence-Aware-Layer-wise-Evolution-Decoding-A564/README.mdanonymous. Benchmarks: https://huggingface.co/datasets/truthfulqa/TruthfulQA and https://huggingface.co/datasets/hotpotqa/HotpotQA.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.