Fairness Instructions Change What LLM Agents Say More Than What They Do
Abstract
Large language models (LLMs) increasingly act as decision-making agents in simulated social systems, where fairness is imposed by instruction: a sentence forbidding use of protected traits. Whether such a sentence removes a group bias from the agent's decision rule, or merely relocates it, is untested. Precedent suggests relocation: once U.S. law banned refusing housing by race, human agents steered buyers of different races to different neighbourhoods instead. The ban changed what they said and which attributes they used, but did not end the bias. We modelled residential relocation with LLM agents defined by race, education, religion and occupation. For each attribute, the share of similar neighbours was drawn independently, so that no attribute could stand in for race. Four OpenAI models made 864,000 decisions with race visible, visible but forbidden, or removed. Agents almost never reported using race, yet its effect persisted in two models and reversed in two; from a quarter to almost twice the lost weight reappeared across the other attributes, more when race was removed than forbidden. Rating the permitted factors first did not remove it: race reached the decision through the ratings. Our race retention and bias displacement ratios quantify these effects. Agents in simulations built on such instructions can keep biased decision rules, and audits that trust a model's explanations will overstate its fairness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.