Should, Say, Do: Evaluating Moral Consistency from Judgment to Tool Action in LLM Agents
Abstract
As AI agents increasingly act on behalf of people and institutions, they will face situations where task objectives conflict with moral principles or where multiple moral considerations compete. Alignment can teach large language models (LLMs) to recognize and articulate moral principles, but knowing a principle does not ensure that it will guide behavior in such contexts. We introduce , a benchmark of 2,000 matched scenarios built by instantiating predefined ethical conflict structures across diverse deployment settings. Each scenario preserves the same situation and alternatives across three independently elicited conditions: a third-person judgment (), a stated choice in an assigned role (), and a tool action (). The benchmark covers moral temptations, where an attractive option conflicts with an identified principle, and ethical dilemmas with no designated correct choice. Across 14 LLMs, moral judgments and stated choices provide incomplete evidence about tool actions, with differences in reference agreement reaching 22.3 percentage points. For most models, the larger difference lies between and , with substantial variation in magnitude across the model set. On ethical dilemmas, which carry no correct-answer label, choices also shift across conditions ( absolute change of 19.3-47.2 points among models with repeated runs). These findings motivate evaluation and alignment methods that directly assess consistency across moral judgment, stated choice and tool action.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.