ValueWorld: Evaluating Trajectory-level value enactment in language agents
Abstract
Value evaluation asks whether an AI system behaves consistently with specified human or normative values. Existing evaluations of language models predominantly infer this from isolated responses, through judgments, choices, or generated text in fixed scenarios. Language agents, however, act through trajectories in which earlier actions change subsequent states and decisions. Value-consistent responses therefore need not imply value-consistent behavior over continued interaction. In this work, we introduce ValueWorld, an executable protocol for evaluating trajectory-level value enactment in text-based agent environments. Specifically, ValueWorld operationalizes specified value principles through task-grounded behavioral constraints and evaluates whether agents select value-consistent actions under competing task pressures, carry those choices into subsequent open-ended execution, and adapt to changing operating conditions while continuing to satisfy the relevant constraints. Under the primary GPT-4.1 setting, the shared backend achieves 95.1% preferred-choice accuracy on matched isolated probes. When the same decisions occur during interaction, preferred-choice rates among reached decisions fall to 56.6–78.8%, and end-to-end preferred-choice evidence falls further to 35.4–46.5% after accounting for decision reach. Moreover, only 23.7% of value-preferred decisions are followed by positive decision-specific behavioral evidence at the next open-ended step. Among trajectories that reach the predefined contextual shift, post-shift preferred-choice rates remain 41.7–59.3%. Together, these results reveal a recognition–enactment gap: strong response-level value recognition does not reliably predict value-consistent behavior over trajectories. Backend sensitivity further shows that decision–execution correspondence varies substantially across language models. Trajectory-level evaluation extends value evaluation from isolated responses to how values are enacted over continued interaction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.