PraxiScene: Benchmarking Agent Social Agency in Generated Open-Goal Interactive Scenarios
Abstract
As AI agents increasingly operate in social environments, they often face situations with no predefined objective and incomplete decision-relevant information. Recent benchmarks have moved toward interactive social behavior, dynamic information, and complex social reasoning, but typically retain a given task, social goal, or inference target. It remains unclear whether social agents retain their performance when deciding what the situation requires and what evidence to seek becomes part of the interaction itself. We formulate this setting as open-goal social agency and introduce PraxiScene, a benchmark combining structured story generation, an interactive environment for action-driven branching narratives, a suite of open-goal social scenes, and trajectory-level evaluation. Its 1248 frozen social situations comprise two story types: one exposes open branching interactions towards different story endings, while the other additionally contains intermediate states with hidden information that reshape subsequent choices. These stories unfold through grounded agent actions and environment responses, producing auditable interaction trajectories. PraxiScene evaluates story completion while characterizing how agents adapt their policies to the situation and available information, and what values are expressed along each trajectory. Matched interventions vary Goal, Information, and Value specification, with humans as a behavioral reference. Across six backbones and two agent systems, supplying Information sharply shortens interaction but reduces CR policy agreement by about 11 percentage points; supplying a Goal instead raises agreement by up to 9 points, while Value guidance selectively increases its targeted realization by 10.8 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.