acceptodds
Under review as a conference paper at ICLR 2027

Fully Private AI Alignment to Human Values

Abstract

Reinforcement learning from human feedback is the preeminent method for aligning powerful AI systems with human preferences and values. Training data for RLHF typically consists of a prompt taken from a conversation with a human along with multiple AI responses, ranked according to the human’s preferences. Data of this form is collected both from professional raters, as well as from everyday users of publicly available AI chatbots, thus presenting a serious privacy risk if used for training. In this paper we study the problem of alignment from human preferences while preventing privacy leaks of human-written training data. We design algorithms that provably achieve -differential privacy in the standard Bradley-Terry and Plackett-Luce models that fully privatize the entire interaction between human and AI. Furthermore, our algorithms achieve asymptotic statistical efficiency that matches the minimax optimal non-private RLHF algorithms in this setting. These results demonstrate that, in a standard theoretical setup, there is no tension between maintaining privacy and efficient alignment training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.