acceptodds
Under review as a conference paper at ICLR 2027

UNVEILING THE REASONING PROCESS OF LARGE LANGUAGE MODELS

Abstract

Layerwise analyses of transformers often identify where information is decodable, yet decodability alone cannot distinguish information that is merely propagated from information that is reorganized, used, or newly composed. We study how relational information develops across depth in six base language models by combining information decomposition, representational geometry, causal interventions, and controlled functional analysis. Across naturalistic task trajectories, interior layers consistently form a transition regime in which residual writes change direction, attention-head synergy becomes enriched, and the occupied representational support expands. Controlled interventions further show that composed relations enter causal use before direct readout, while exact functional decompositions place the stable emergence of higher-order relations later in the network. Rank-guided head ablations connect this organization to behavior and reveal that its functional expression depends on the task. Together, these results support a view of transformer depth as progressive relational construction: aligned low-order structure is reorganized into causally effective compositions, which subsequently stabilize and align with the output space.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.