acceptodds
Under review as a conference paper at ICLR 2027

Right on Average, Wrong per Respondent: A Mechanistic Diagnosis of LLM-Simulated Choice

Abstract

Language models are increasingly used as synthetic survey respondents, and what those applications need is how a choice moves when an attribute changes, not how plausible a single answer looks. Existing evaluations score the level of an answer and not its response to a change. We introduce MagBench, which perturbs one attribute of a choice situation at two sizes, the smaller half the larger, on five datasets of real human choices, and asks whether the response scales in proportion. In aggregate it does. For an individual respondent it does not, and most of that respondent's movement has nothing to do with the size of the change. The aggregate is right because those respondent-specific parts cancel, and what cancels is not noise but a displacement that reappears at both sizes of change and under a different prompt. Inside the models the size of the change is decodable and reaches the subspace that carries the response, and neither fact predicts behavior. Writing the displacement one respondent's perturbation produces into another respondent's run moves the response to the donor and not the recipient, at 0.89 against over 67 runs and five model families. It holds even in a model whose choices follow the human direction 51 percent of the time, so whether this structure is present and whether the behavior it should drive is valid are separate questions. Encoding a quantity is not using it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.