acceptodds
Under review as a conference paper at ICLR 2027

Do Language Models Dream of Answers before They Speak?

Abstract

At the core of language model intelligence lies the formation of task-relevant representations, internal encodings of the information underlying the model’s inferential outcomes. However, the representational dynamics governing the transformation of latent states into model outputs remain insufficiently characterized. We study the evolution dynamics of task-relevant representations during inference and observe task-relevant direction supporting highly accurate readout that generalizes across datasets emerges abruptly at an intermediate model layer. To characterize the geometry underlying this readout, we develop and empirically validate a theoretical framework grounded in real-world language models, accounting for interactions among components and deriving empirically verifiable criteria for the direction’s spectral dominance. Furthermore, we substantiate the readout’s effectiveness and semantic alignment with the model’s native answer preferences through visual evidence reranking experiments for retrieval-augmented generation. By connecting the internal organization of task-relevant representations to model outputs, our work provides a principled basis for interpreting representational dynamics and applying internal model judgments in real-world systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.