acceptodds
Under review as a conference paper at ICLR 2027

E-CALR: Evidential Conflict-Aware Layer-wise Logit Refinement for Factuality-Oriented Decoding

Abstract

Large language models (LLMs) can generate fluent but factually incorrect outputs. Layer-wise decoding seeks to improve factuality by exploiting predictions exposed at intermediate Transformer layers. A representative method, SLED, uses directional alignment in cross-layer logit evolution to construct intermediate-layer estimates for refining the final prediction. However, directional alignment alone does not reveal whether an intermediate prediction expresses a clear preference among specific token alternatives. A diffuse prediction may therefore receive a high alignment score despite providing weak token-specific evidence. This motivates treating intermediate predictions as potential sources of interference with the final prediction, whose influence should depend on how strongly they concentrate on specific token alternatives. We refer to this degree of concentration as predictive commitment. To this end, we propose Evidential Conflict-Aware Layer-Wise Logit Refinement (E-CALR), a training-free decoding framework that identifies cross-layer conflict backed by committed evidence and uses it to selectively suppress interfering intermediate-layer information. E-CALR represents each layer prediction in terms of token-specific committed evidence and nonspecific ignorance, measures evidential conflict with the final-layer prediction, and further modulates this conflict by predictive commitment so that diffuse intermediate predictions exert less influence. The resulting weights aggregate intermediate-layer distributions into a token-level suppression signal that iteratively refines the final-layer logits while preserving the original final-layer top- candidate set. Across nine LLMs, E-CALR improves factual candidate ranking over representative layer-wise decoding baselines and also generalizes to autoregressive reasoning generation. Compared with SLED, E-CALR improves model-averaged performance across all reported factual-ranking metrics and achieves higher accuracy for all nine models on both GSM8K and StrategyQA.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.