acceptodds
Under review as a conference paper at ICLR 2027

Minority Welfare Under LLM-Governed Allocation: A Systematic Empirical and Mechanistic Study

Abstract

Large language models are increasingly proposed as planners that turn public preferences into resource allocations. Such a system can acknowledge competing views in its explanation while placing a disproportionate share of allocation error on people whose preferences differ from the majority. We test for this failure in a controlled negotiation where agents (LLM-simulated participants) may strategically misreport their preferences and the planner observes only part of the population each round. We compute a Nash Social Welfare optimum from each agent's true, hidden preferences and measure how the resulting welfare loss is distributed between minority and majority agents. Across four closed models, a minority comprising ten percent of the population bears more than four times its population share of welfare loss. This disparity remains under changes to preferences, agent behavior, agent capability, planner instructions, and population size. On two open-weight planners, recognition of preference misreporting is decodable from late-layer activations before allocation, yet steering this signal does not reliably improve welfare. We call this pattern a crystallized prior, in which the planner registers preference misreporting while allocation commitment remains effectively insensitive to it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.