acceptodds
Under review as a conference paper at ICLR 2027

Test-Time Personalization via DPO with Latent Reward Features

Abstract

Standard preference alignment of language models (LMs) relies on a single reward model (RM), overlooking the diversity of human preferences. While recent work on personalized RMs address this by learning reward features—representing shared latent preference dimensions—together with user-specific weights, updating the alignment objective to these personalized weights may still require repeatedly performing costly preference optimization for each user. In this paper, we propose TPD (Test-time Personalized Decoding), a framework for on-the-fly personalization of the LM that avoids model retraining for each user. TPD first learns a set of shared feature-level policies through a DPO-like objective, bypassing the explicit learning of reward features. At test-time, TPD then balances these policy logits using weights adapted to the user's preferences, following a multi-objective decoding approach. Importantly, we show that TPD naturally extends to a broad class of alignment objectives based on -divergences. Experiments on synthetic and real-world datasets demonstrate effective steering of generation toward maximizing individual user satisfaction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.