acceptodds
Under review as a conference paper at ICLR 2027

Preference Stealing from Best-of-: The Privacy Risk of Inference-Time Alignment

Abstract

Inference-time alignment methods are gaining popularity since they improve the helpfulness and safety of language models' responses without training. Most methods rely on a reward model (RM) to raise the probability of good responses. The RM encodes preferences from expensive, proprietary, and often sensitive human-annotated training data. Both the RM itself and these preferences are therefore valuable targets for attackers. However, prior work on preference privacy has focused on training time, leaving the risk at inference time unexplored. We first raise the concern that inference-time alignment methods might leak the RM's preferences through the responses they serve. We take Best-of- (BoN) sampling, one of the most widely used methods, as the representative case. BoN keeps its RM in the serving loop, so every served response implicitly reveals a preference signal from the RM. We study how attackers might acquire the RM's preferences under different deployment scenarios, which expose different amounts of information about BoN's alignment process. Our experiments show that even under the strictest scenario in which BoN serves only the final response, attackers can still recover the preferences behind the RM's training data with about accuracy, against for chance. Finally, we take a first step toward a defense through output perturbation, yet the attack survives almost intact. BoN keeps leaking the RM's preferences, and closing the leak remains an open problem.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.