PerSPL: Personalized Semi-supervised Preference Learning
Abstract
Personalized reward models require user-specific preference comparisons, but collecting such feedback is costly and provides only limited supervision for each user. At the same time, large pools of responses can be generated without user identities or preference labels, yet existing personalized reward models leave this data unused. Directly learning from these responses is difficult because pairwise objectives constrain relative scores within a prompt, whereas pseudo-labeling individual responses requires scores that are comparable across prompts for the same user. We propose PerSPL (Personalized Semi-supervised Preference Learning), a backbone-agnostic framework that uses pair-mean score centering to stabilize this cross-prompt score coordinate and iteratively turns unlabeled responses into user-specific training signals. Across three primary benchmarks with a full suite of baselines, including Reddit TLDR with real-user preference comparisons, PerSPL achieves the highest mean accuracy for both seen and unseen users. A focused evaluation on Chatbot Arena further shows gains over the two strongest personalized baselines from the primary suite on real-user open-domain preferences. These results show that unlabeled responses can provide additional supervision for personalized reward modeling when per-user preference labels are scarce.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.