Deep reinforcement learning agents represent reward-dependent belief geometry
Abstract
Bayesian beliefs are sufficient statistics of the history of a partially observed process, and classical separation results show that optimal control can be attained using these beliefs. However, it is unclear whether belief-like representations actually emerge in model-free deep reinforcement learning (RL), which trains policies not for prediction but to maximize reward. To address this question, we study transformer policies trained by policy gradients and fit affine probes into their residual streams, scanning for Bayesian posteriors. Experiments show that reward alone suffices to recover beliefs, even in the absence of latent-state supervision or auxiliary losses. However, we also find that the full belief geometry is not always what the agent learns. Using a family of tasks, in which the reward–action structure requires progressively finer distinctions between latent states, we show that distinctions that are irrelevant for control and reward are selectively lost, and that the learned representation is better described by the beliefs of a coarse-grained POMDP than by the full environmental posterior. Changing the reward leads to the same effect: only reward-relevant latent factors are accounted for within the belief geometry. Taken together, these results suggest a learning-dependent version of the separation principle, in which the environment determines how beliefs are updated while the reward determines which beliefs are worth maintaining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.