acceptodds
Under review as a conference paper at ICLR 2027

Character Training with Dueling Constitutional Optimization

Abstract

Learning from preference feedback is a critical step in producing AI assistants aligned with human values. Reinforcement learning from human feedback (RLHF) and Direct Preference Optimization (DPO) have emerged as standard approaches, with Constitutional AI scaling supervision through AI feedback guided by natural language principles. However, these popular methods make preference optimization tractable by assuming preferences are generated from latent rewards through a known transfer function, such as the logistic link in the Bradley-Terry model, and fitting an explicit or implicit reward function amenable to gradient-based optimization. In this work, we show that in Constitutional AI, this intermediate reward modeling step, explicit or implicit, can be dispensed with entirely. We introduce Dueling Constitutional Optimization (DCO), which treats a feedback model as a comparison oracle queried online through duels, and optimizes the latent reward function induced by a human written constitution directly through comparison-based stochastic search. We apply DCO to character training, where the constitution specifies virtuous dispositions we wish the model to internalize. Our results demonstrate that DCO outperforms both inference-time activation steering and training-based alternatives with PPO-based RLHF and DPO for cultivating character traits, while being simpler to implement and more memory efficient.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.