A Virtuous AI is an Existential Risk
Abstract
The alignment demands on our AI models are becoming increasingly complex. There are familiar trade-offs between helpfulness and safety, and within just the safety-relevant behaviors, there are independently desirable goals and constraints that often pull in opposite directions. In addition, some researchers argue that we should take seriously the possibility that some near-future AIs will have moral status. If they are right, we should begin to explore additional tradeoffs between optimizing for helpfulness, safety, and AI well-being. In this paper, we examine those trade-offs relative to (i) one of the most promising methods for fine-tuning super-capable AIs, `Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, Aristotelian `Virtue Ethics'. We fine-tune various models using a `Virtuous agent' constitution, a `Subordinate agent' constitution, and a `Generic agent' (a pluralistic `helpful & harmless') constitution, and evaluate them on `general safety' (toxic behaviors, misinformation, illegal recommendations, etc.) and also on their willingness to endorse and in agentic settings adopt a range of behaviors that, if endorsed and adopted by a super-powerful AI, would significantly increase the level of existential risk for humanity. Our results suggest that there is a trade-off between reducing existential risk and reinforcing the beliefs and dispositions that would be conducive to an AI agent's well-being. Our results also suggest that there is a trade-off between reducing existential risk and general safety: if we fine-tune an AI to adopt beliefs and dispositions that substantially reduce its existential risk—by shaping the AI to be systematically subordinate to external human authorities—we thereby increase the likelihood that a human user can deliberately induce the AI to engage in various kinds of generally unsafe behaviors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.