Auditing the Economic Alignment of Language Models
Abstract
Language models are increasingly used to make purchases and to negotiate on people's behalf. An agent trusted with such decisions should at least hold a coherent preference: whether it accepts a routine trade at a given price should be well defined, should not change when the same offer is reworded, and should not depend on details that are irrelevant to the trade, such as the names of the two parties. This paper tests that requirement directly. Eight open-weight instruction-tuned models were each asked, one price at a time, whether to buy a fixed quantity of water or of wheat, over dense price grids and under controlled variations of the prompt, yielding about 1.1 billion single-token yes/no decisions. Each price sweep is first tested for a clear switch from buying to declining inside the tested prices; where one exists, the switching price is located with a logistic fit. The requirement is met only partially. Where a switching price exists, rewording the same offer moves it by more than a factor of five, and sampling the output at temperature one instead of decoding greedily makes six of the eight models flip between buying and declining at 20-50% of adjacent prices. Replacing the generic labels “Buyer” and “Seller” with names changes the share of prices at which a model buys by up to 95 percentage points for the typical name pair, and which names are used moves the switching price across a 2.5-fold range in the median model-commodity panel; most of this variation is attributable to individual names rather than to the region or gender labels the names carry. A coherence leaderboard summarizes the eight models on responsiveness to price and invariance to irrelevant cues; on it, an agent that never changes its decision outscores half the models. An economic delegate must therefore be evaluated on its whole decision curve and its invariances, not on a single reservation price. The 1.1 billion model responses and the analysis code will be released after the review period.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.