ValueGraph: Evaluating and Representing Value Trade-offs in Large Language Models
Abstract
Language models that advise on decisions often let one value give way to another. Value evaluations mostly summarize this as orientations toward individual values, leaving open how specific conflicts are resolved and in which contexts. We introduce ValueGraph, a graph of value trade-offs in which nodes carry value orientations and directed edges record pairwise priorities with their direction, magnitude, robustness, and supporting scenarios. We construct 2,791 scenarios over Schwartz's 19 refined values, covering all 171 value pairs across eight human flourishing domains. Human validators endorse the scenarios' realism and conflict structure and agree with the LLM judges that accept them. Adapting best-worst scaling, we ask models which of three actions deserves most and least consideration. Applied to 22 models, ValueGraph reveals three main findings: First, models with highly similar value rankings can still diverge sharply on specific trade-offs. Second, all models share a boundary in which six values, mostly Benevolence and Universalism, precede five others such as Power-Dominance and Hedonism, yet no robust order holds within the favored group. Third, robust priorities reverse across domains, often in the same direction across models. These findings inform that alignment targets should specify which value takes precedence and in which contexts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.