acceptodds
Under review as a conference paper at ICLR 2027

Which State Should Prediction Read? A Mechanistic Analysis of State-Prediction Separation

Abstract

State-Prediction Separation (SPS) improves language modeling with separate persistent-state and next-token prediction slots, but sharing all weights leaves it unclear whether and how their roles specialize. We trace information flow after giving the streams separate weights. State concentrates predecessor information, while prediction shows stronger induction matching, retrieving continuations of earlier token occurrences. Ablating state's previous-token heads selectively weakens prediction's retrieval, and restricting access to older states worsens loss. Restricting prediction to the deepest state levels costs far less than restricting it to the shallowest, yet SPS never reads its top state layer. We therefore give every prediction layer the final state (a final-state read), a model we call Sequential. With each tower as deep as the Transformer, Sequential lowers seed-mean loss by 0.031 nats/token versus the same-depth read (Two-tower) at matched parameters and floating-point operations (FLOPs). A smaller Sequential model with half the layers per tower comes within 0.005 nats/token of SPS at 55% of its FLOPs and 0.064 below the Transformer at matched non-embedding parameters and FLOPs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.