A Steerable Internal Direction Links Distinct Mental-Accounting Biases in Large Language Models
Abstract
Large language models (LLMs) increasingly stand in for human participants in economic and behavioral research, and reproduce many human decision-making biases in their answers. But matching an answer does not reveal whether a model carries the underlying concept or only imitates familiar patterns, which determines how far such synthetic respondents can be trusted. We probe this using mental accounting, the tendency to treat money as non-fungible, valued differently by its source or use, which behavioral economists have cataloged as several separate biases: segregating gains, integrating losses, the sunk-cost effect, and transaction utility. Rather than score answers alone, we read and edit the models' internal activations, the numerical state a model builds while processing a prompt. In LLaMA-3-70B, the three categorical biases share a common internal direction: reading it from the model and then adding or subtracting it shifts economic choices as predicted, while a matched random direction does not, showing the representation is causal, not merely correlated with behavior. A related transaction-utility direction widens the venue-dependent willingness-to-pay gap. A controllable de-biasing direction recurs across independently built large models (Qwen2.5-72B, Gemma-2-27B) but is largely absent in three smaller ones, which still show the biases yet offer little steerable control. Controllability is more evident at larger scale in the models tested. The results reveal a mechanism for why high-capacity LLMs replicate economic behavior, and a standard for auditing them as synthetic respondents: match not only their choices, but whether the structure behind those choices is stable and controllable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.