acceptodds
Under review as a conference paper at ICLR 2027

Reading Between the Logits: Uncovering Vocabulary-Grounded Belief in LLMs

Abstract

Understanding the internal representations guiding an LLM's judgement, particularly when these diverge from its verbalised responses, is important for interpretability, reliability and safety. Existing approaches typically extract such signals using supervised probes, which require labelled data and training for each new setting. Lens methods, such as the logit lens and J-lens, instead project hidden states into vocabulary space, providing readouts of intermediate model representations. We introduce an unsupervised method that clusters vocabulary-level patterns in lens readouts to recover concepts: interpretable groups of tokens that express the LLM's judgement of a factual claim, such as whether it treats the claim as true, false or uncertain. These concepts generalise across prompt styles and model families, and exchanging their directions flips the model's verdict across claims. From these concepts, we compute a training-free signal that approaches the performance of supervised linear probes and remains comparatively stable when a user's conflicting belief alters the model's verbalised judgement. Finally, a complementary topic modelling analysis reveals readout patterns related to truth, falsity, misconception and contested claims, and how these develop across layers. These patterns are already present in base models, but become aligned with verbalised judgements only after instruction tuning. Together, concepts and readout patterns give an unsupervised, interpretable and inexpensive account of the representations behind a model's judgements, and of when these diverge from what it verbalises.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.