LLMs Contain Multitudes: How Deployment Context Reshapes Model Preferences & Values
Abstract
The preferences and values attributed to large language models (LLMs) shape judgements about their bias, alignment and suitability for deployment. Recent pairwise evaluations increasingly ascribe stable, model-level preference and value systems to LLMs. However, accompanying robustness checks are limited to incidental prompt perturbations such as syntax variation and option reordering. They do not test whether the measured properties survive when the surrounding task context changes, as is common in real deployments. An audit may therefore provide false reassurance about the biases most consequential for a model's intended use. We test this context dependence directly using two established pairwise paradigms: country preference ranking and utility elicitation. In both, we make *deployment context* – the high-level task the model performs while making concrete value-dependent choices – our controlled variable. Across five LLMs and over 1.2M pairwise elicitations, we vary deployment contexts across framings such as writing a Reddit post or a news article. In preference rankings over 15 countries, context induces widespread, statistically significant rank shifts; each model's bias shifts systematically across contexts. Well-accepted patterns in prior work, such as Global North favouritism, also appear context-dependent. In utility elicitation over 50 outcomes, broad ordering remains consistent, but rankings within domains vary substantially, and cardinal utility ratios between outcomes (e.g. the utility of averting a death relative to receiving $1M) shift by at the median. Reported model-level preferences and utilities are therefore better understood as context-conditioned measurements than fixed model-level properties: bias and alignment conclusions drawn under one framing need not generalise to another.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.