acceptodds
Under review as a conference paper at ICLR 2027

Emergent alignment and the projectability of ethical personas

Abstract

Recent work on "emergent misalignment" has shown that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the "persona selection" model (PSM) hypothesis that, during pre-training, LLMs learn to simulate many different characters and perspectives, which can then be elicited and refined during post-training. Inspired by those results, this paper investigates the converse phenomenon, "emergent alignment", and uses it to support and refine the PSM and motivate a novel desideratum for alignment. We finetune a helpful-only model on broad and narrow safety tasks. To create SFT samples, we follow the "Constitutional AI" (CAI) approach and use four constitutions that could be part of reasonable alignment strategies: three drawn from canonical ethical systems (deontology, consequentialism, and virtue ethics) and one corresponding to the strategy of having AIs be subordinate to, and concerned solely with, the good of humanity. Each of the four constitutions is used to train a corresponding model, and for each of these models we show that fine-tuning on two narrow safety sub-categories (harassment and illegal behaviors) reliably induces emergent alignment. Specifically, our narrowly aligned models perform significantly better than the helpful-only source model on a benchmark covering a representative sample of general safety categories, as well as on specific safety categories that were carefully filtered out of the data sets used for narrow alignment finetuning. To test the PSM, we also use a fine-grained multidimensional persona-diagnostic which includes dimensions for deontological, consequentialist, virtue-ethical, and "subordinate" personas. For each constitutionally finetuned (broad and narrow) model, we evaluate how well their behavior matches their expected signature profile (given their anchor constitution). Our results show that our CAI models acquire their expected personas—e.g., the model narrowly fine-tuned on SFT samples created using the consequentialist constitution agrees significantly more with consequentialist than deontological beliefs. At the same time, both our coarse and fine-grained evaluations show that there are significant differences across our (broad and narrow finetuned) CAI models in how well they project. Based on those results, we argue that alignment strategies should be evaluated, not just on their (in-distribution) general safety performance, but also specifically on their degree of projectability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.