acceptodds
Under review as a conference paper at ICLR 2027

Transformers develop separable posterior and value representations in a simple decision task

Abstract

How does a transformer make decisions when it is uncertain about the world? A network could form a posterior over hidden states and value its options under that posterior, or it could map observations to choices with no such intermediate structure. We investigate this question with a transformer trained by reward alone on a simple decision task. Hidden states are assigned payoffs, and offers composed of sets of hidden states are presented to the transformer. The offers and state payoffs are given in context, and the exact posterior and the value of every offer are known in closed form. We anticipate and then find that both quantities are linearly represented in the residual stream (), in separate subspaces, and that the representations extend to offers never seen in training. Interventions show that both are used for choice: steering the posterior shifts choice by the amount the ground truth predicts, and does so through the value subspace, while steering the value shifts choice and leaves the posterior fixed. None of this is required by the objective, which constrains only the difference between the offer values. A pretrained language model, Qwen3-1.7B-Base, trained on the same task, develops the same organization. Our work extends the posterior geometry found under next-token prediction to choice, and gives a setting in which the quantities that drive a model's decisions can be measured against ground truth and intervened on.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.