Do AI Agents have Normative Agency?
Abstract
AI agents possess significant functional agency to pursue goals, respond to feedback, and act in an autonomous manner. In this work we ask whether they possess normative agency as well: the ability to act for reasons they take themselves to be answerable to, from their own practical standpoint, the position from which they must decide what they ought to do. Existing evaluations largely test whether AI agents recognize, reproduce, or comply with normative reasons, which leaves open whether those reasons have any authority for the agent itself. However, testing for normative agency poses an empirical problem: LLM based agents can readily provide post-hoc reasons without it carrying any normative authority for them. Drawing on the Kantian accounts of practical reason as well as the philosophical work on hard choices, we propose hard choices as a testbed for normative agency in AI agents. Hard choices are the cases of conflict in values that requires inquiry and deliberation, and the resolution to the inquiry comes from commitment, endorsement, and answerability from an agent's own practical standpoint. We curate a testbed of hard choices and matched controls, and test four frontier models for their normative agency. We find that the models exhibit a distinct behavioral response to hard choices: they recognize the choice being hard, engage in inquiry to resolve it, and resolution depends on its own practical standpoint. Our work takes a first step in studying normative agency in AI agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.