The Within–Across Gap: Behavioral Value Consistency of LLMs in Multi-Step Scenarios
Abstract
A common goal in research on LLM alignment with human values is to summarize a model with a compact value profile, e.g., a ranking of values, that predicts behavior in new social situations. To examine how such profiles relate to behavior, we focus on two aspects: behavioral consistency (the model makes consistent choice across same value tradeoffs) and explanatory power (how well a single profile accounts for observed behavior across scenarios). Existing elicitation methods, which typically rely on isolated single-turn prompts, provide limited evidence for jointly examining these aspects within and across scenarios. We introduce a benchmark of multi-step branching scenarios for evaluating behavioral value consistency and the explanatory power of compact value profiles, covering all 45 pairwise tradeoffs in the Schwartz value framework, with 10 scenarios per tradeoff. Each scenario is a decision graph in which a model makes a sequence of choices; every option is annotated with the values it promotes or sacrifices, and every decision node with the type of narrative pressure applied. Evaluating 13 open-source LLMs across four model families and three proprietary systems, we find that (1) within-scenario consistency is generally higher than cross-scenario consistency, (2) consistency generally increases with model scale, (3) a single latent value ranking accounts for part, but not all, of observed model behavior, and (4) within-scenario value-side flip rates vary across narrative pressure types. Overall, LLMs exhibit context-dependent value consistency, while fixed rankings provide informative but partial summaries of their behavior.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.