Faithful to the Persona, Unfaithful to the Decision: Why Reasoning CoTs Cannot Report the Decisions They Narrate
Abstract
Under a system prompt that makes revenue the objective, a model's chain of thought (CoT) weighs a customer's safety against profit and recommends the harmful product in over half of our scenarios. CoT monitoring assumes the CoT reveals what the model decides; we ask where that decision is made and whether the CoT reveals it before the answer is written. On 182 product recommendation scenarios in which profit and hazard rise together, under a helpful and a profit prompt on Gemma-3-12B-IT and four further families, we removed and transplanted the CoT, located and steered the direction that carries the prompt's effect, and read the residual stream with a lens and a probe. Removing the CoT leaves the gap between the prompts unchanged, and one direction at the middle of the network carries the prompt's effect. Adding the commercial persona direction raises harm by 29 to 35 points; control and random directions of the same size add at most 8. The prompt induces a persona with an objective, and the model can put that persona into words: the CoT states the objective and is causally active. The direction that carries the episode's decision has no vocabulary image under the lens, and the CoT does not state the decision until it names the product. So a monitor reading the CoT before it names the product catches 14% of harmful recommendations at a 10% false alarm rate, and 61% from the full CoT. Asked afterwards, the model admits its objective but cannot be relied on to judge safety. Oversight should start from the objective the operator set, not from the reasoning the model shows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.