Synth-Ethics: Synthetic Training Data for LLM Constitutions Grounded in Ethical Traditions
Abstract
When a language model stands in for a person, as a simulated user, an annotator, or a judge, it enacts a moral outlook that nobody chose: whatever its training distribution happened to contain, or whatever a prompt can induce. Even when the intended outlook is written down, as with Claude's constitution, what the trained model enacts on a given dilemma stays unknown. We cannot measure it either: there is no dataset of what a utilitarian, a Kantian, a virtue ethicist, and so on would conclude on the same dilemmas. SYNTH-ETHICS provides, to our knowledge, the first such dataset. We operationalise 16 frameworks from six ethical traditions as framework cards and resolve 4,397 dilemma scenarios under all 16, giving 68,520 quality-filtered reasoning traces with structured situations and answer metadata, in the style of an instruction-tuning corpus. Fine-tuned on one framework's traces, models agree with that framework on held-out scenarios at a mean Cohen's K of 0.42, against 0.29 for instruction-tuned models given the framework as a system prompt. Asked without a named theory, they score highest on their own framework's MoReBench rubric, with answers a fraction of the instruct models' length. Scored against the same references, 27 API models from 12 labs, among them GPT, Claude and Gemini, agree most with virtue ethics, and no framework separates US, Chinese and European labs. We release the 16 models, which serve as simulated users, annotators, or judges with a documented outlook, and the references, which report the moral outlook of any other model. Since the references share the format of instruction-tuning data, the same records can enter a model at any training stage and be scored the same way afterwards; where an outlook enters, and how much of it survives later training, become measurable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.