acceptodds
Under review as a conference paper at ICLR 2027

Candidate-Conditioned Representations for Factual Reliability in Language Models

Abstract

A language model’s confidence in an answer does not necessarily reflect whether that answer is factually correct, exposing a gap between output confidence and candidate-level factual reliability. We investigate whether the internal representation induced by a specific candidate answer provides a distinct signal of candidate-level factual reliability beyond the information available from the model’s output distribution. We construct a controlled candidate-conditioned evaluation framework that holds the underlying prompt and competing candidate set fixed while comparing internal representation signals with candidate probability, likelihood rank, token-level uncertainty, entropy, and statistics of the model’s output distribution. Across controlled experiments, candidate-conditioned representations exhibit measurable and reproducible differences between reliable and unreliable candidates, with the observed signal varying across representation levels and evaluation conditions. Our results show that factual reliability is reflected in candidate-conditioned internal states in ways that are not fully captured by conventional output-distribution measures, establishing candidate-conditioned representations as a distinct empirical source for factual-reliability assessment, and show that candidate reliability is reflected not only in the model’s final output probabilities but also in the internal state induced by the candidate itself.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.