Reinforcement Learning in Language Models Recruits an Evaluative Axis
Abstract
How does reinforcement learning shape a language model’s internal representations? We present evidence that RL recruits a pre-existing evaluative axis: a direction associated with success and failure that broadly modulates model behavior. We train several language models in a novel, semantically neutral maze environment, extract concept vectors for rewarded and punished trajectories, and evaluate those vectors on tasks unrelated to the maze. The punishment vector promotes failure and impossibility tokens, aligns with negative emotion concepts, and, under steering, induces negative self-reports, pathological backtracking, refusal, and uncertainty. The positive reward vector behaves as the mirror image, and the two are nearly antiparallel. These effects are robust across training and extraction seeds, model families, and environmental controls, and largely persist when we replace RL with supervised fine-tuning. Importantly, these effects appear in the models before any maze training. This supports recruitment of pre-existing evaluative structure, rather than creation of the behavioral effects from scratch. Additional environments yield partially aligned directions with similarly signed steering effects. Together, these results demonstrate that minimal reward signals can recruit directions that broadly affect model behavior, with implications for interpretability, post-training dynamics, and alignment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.