Human Non-Choice in LLMs: Internal Support, Limited Expression
Abstract
Large language models (LLMs) are increasingly used to simulate human responses under demographic and contextual personas, yet matching directional opinions alone may overlook an important aspect of human response behavior: people sometimes decline to commit to an available position. We study whether persona-conditioned LLMs systematically underexpress “Can’t choose” (CC) responses and whether this mismatch can be mitigated through activation-level interventions. Using 42 survey items spanning environment, family, and health across six countries, we construct cell-balanced human targets and contrast them with LLM response probabilities. We derive shared activation directions from teacher-forced contrasts between CC, midpoint, and directional responses, and intervene at inference time using three complementary mechanisms: component-wise head gain, angular rotation in residual space, and additive steering. Intervention strengths are selected to minimize divergence from human CC rates, while midpoint behavior and the full six-category response distribution remain independent evaluation outcomes rather than optimization targets. We further compare learned interventions with matched random-direction controls and a direct semantic-logit bias, and evaluate generalization on held-out respondents under multiple answer-order permutations. This framework separates the problem of behavioral fidelity from simply increasing indecision, providing a controlled test of whether internal representation interventions can correct systematic omissions in persona-conditioned LLM simulations without retraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.