acceptodds
Under review as a conference paper at ICLR 2027

Human Non-Choice in LLMs: Internal Support, Limited Expression

Abstract

Large language models (LLMs) are increasingly used to simulate human responses under demographic and contextual personas, yet matching directional opinions alone may overlook an important aspect of human response behavior: people sometimes decline to commit to an available position. We study whether persona-conditioned LLMs systematically underexpress “Can’t choose” (CC) responses and whether this mismatch can be mitigated through activation-level interventions. Using 42 survey items spanning environment, family, and health across six countries, we construct cell-balanced human targets and contrast them with LLM response probabilities. We derive shared activation directions from teacher-forced contrasts between CC, midpoint, and directional responses, and intervene at inference time using three complementary mechanisms: component-wise head gain, angular rotation in residual space, and additive steering. Intervention strengths are selected to minimize divergence from human CC rates, while midpoint behavior and the full six-category response distribution remain independent evaluation outcomes rather than optimization targets. We further compare learned interventions with matched random-direction controls and a direct semantic-logit bias, and evaluate generalization on held-out respondents under multiple answer-order permutations. This framework separates the problem of behavioral fidelity from simply increasing indecision, providing a controlled test of whether internal representation interventions can correct systematic omissions in persona-conditioned LLM simulations without retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.