Character Training with Dueling Constitutional Optimization
Abstract
Learning from preference feedback is a critical step in producing AI assistants aligned with human values. Reinforcement learning from human feedback (RLHF) and Direct Preference Optimization (DPO) have emerged as standard approaches, with Constitutional AI scaling supervision through AI feedback guided by natural language principles. However, these popular methods make preference optimization tractable by assuming preferences are generated from latent rewards through a known transfer function, such as the logistic link in the Bradley-Terry model, and fitting an explicit or implicit reward function amenable to gradient-based optimization. In this work, we show that in Constitutional AI, this intermediate reward modeling step, explicit or implicit, can be dispensed with entirely. We introduce Dueling Constitutional Optimization (DCO), which treats a feedback model as a comparison oracle queried online through duels, and optimizes the latent reward function induced by a human written constitution directly through comparison-based stochastic search. We apply DCO to character training, where the constitution specifies virtuous dispositions we wish the model to internalize. Our results demonstrate that DCO outperforms both inference-time activation steering and training-based alternatives with PPO-based RLHF and DPO for cultivating character traits, while being simpler to implement and more memory efficient.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.