acceptodds
Under review as a conference paper at ICLR 2027

Behavior Underspecification in Large Human Language Models

Abstract

Large human language models (LHLMs), language models designed to simulate human responses, are increasingly used as scalable simulators of human behavior. Given a description of an individual, an LHLM generates responses from that person’s perspective, enabling applications in behavioral science, marketing, public policy, and beyond. But what kind of human behavior do they actually model? Behavioral science has long documented a consequential gap between what people say they do and what they actually do. We hypothesize that LHLMs may not reliably distinguish between these constructs because their training data disproportionately capture what people say, while their training objectives do not explicitly differentiate among them. We refer to this ambiguity as behavior underspecification. To study it, we evaluate ten contemporary LHLMs on a cross-domain dataset comprising 98 matched measures of what people said and eventually did, derived from more than 37K human responses. With experiments, we find that behavior underspecification is widespread, that LHLMs model what people say better than what people do, and predictions of what people do often reflect what people say. The effect of behavior underspecification persists even in otherwise accurate models. Fine-tuning on human-response data does not necessarily resolve this problem and can sometimes exacerbate it. Together, our findings show that evaluating LHLMs requires asking not only how accurately they simulate human responses, but precisely what aspects of human behavior they simulate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.