Fully Private AI Alignment to Human Values
Abstract
Reinforcement learning from human feedback is the preeminent method for aligning powerful AI systems with human preferences and values. Training data for RLHF typically consists of a prompt taken from a conversation with a human along with multiple AI responses, ranked according to the human’s preferences. Data of this form is collected both from professional raters, as well as from everyday users of publicly available AI chatbots, thus presenting a serious privacy risk if used for training. In this paper we study the problem of alignment from human preferences while preventing privacy leaks of human-written training data. We design algorithms that provably achieve -differential privacy in the standard Bradley-Terry and Plackett-Luce models that fully privatize the entire interaction between human and AI. Furthermore, our algorithms achieve asymptotic statistical efficiency that matches the minimax optimal non-private RLHF algorithms in this setting. These results demonstrate that, in a standard theoretical setup, there is no tension between maintaining privacy and efficient alignment training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.