acceptodds
Under review as a conference paper at ICLR 2027

PerSPL: Personalized Semi-supervised Preference Learning

Abstract

Personalized reward models require user-specific preference comparisons, but collecting such feedback is costly and provides only limited supervision for each user. At the same time, large pools of responses can be generated without user identities or preference labels, yet existing personalized reward models leave this data unused. Directly learning from these responses is difficult because pairwise objectives constrain relative scores within a prompt, whereas pseudo-labeling individual responses requires scores that are comparable across prompts for the same user. We propose PerSPL (Personalized Semi-supervised Preference Learning), a backbone-agnostic framework that uses pair-mean score centering to stabilize this cross-prompt score coordinate and iteratively turns unlabeled responses into user-specific training signals. Across three primary benchmarks with a full suite of baselines, including Reddit TLDR with real-user preference comparisons, PerSPL achieves the highest mean accuracy for both seen and unseen users. A focused evaluation on Chatbot Arena further shows gains over the two strongest personalized baselines from the primary suite on real-user open-domain preferences. These results show that unlabeled responses can provide additional supervision for personalized reward modeling when per-user preference labels are scarce.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.