Detecting Preference Change Is Not Enough: Rethinking Reward Model Personalization
Abstract
Personalized Reward Models (PRMs) train on users' past choices to predict future preferences. Temporal personalization often assumes that once preference change is detected, the model should adapt its use of that history. We argue that this collapses three distinct questions: *is there enough evidence to adapt this particular user? If so, how should memory change? And is the resulting adaptation genuinely driven by that user rather than by population-level regularities?* We study these questions prospectively, using only information available before evaluation. We fix a pairwise PRM with one component learned across training users and vary only how a second, user-memory component constructs the specific user's history. We compare four memory modes: *retain* keeps all past choices, *discount* progressively downweights older choices, *reset* rebuilds recent memory after a detected change, and a multiscale mode learns a mixture of causal memories operating at different timescales. We evaluate these modes on controlled synthetic histories, PENS, PRISM, and Chatbot Arena, while testing whether temporal structure is real, individually actionable, and prospectively useful. We uncover an *information paradox*: preference change can be clearly measurable across users, while an individual user's earlier history does not reliably identify, before seeing the future, which evaluated memory rule will work best. Choosing the best rule *after seeing the outcomes* appears to gain 4 to 6 percentage points, but about half of that is selection optimism, and the gain disappears when the rule must be chosen earlier. We then uncover an *intervention paradox*. A detected user-specific change does not determine the right memory response. Under abrupt change and identical alarms, event-triggered reset beats a temporary fast discount by about 1 percentage point, but which response wins depends on how late the alarm arrives and how far back reset rebuilds memory. Finally, exploratory analysis reveals an *illusion of personalization*. We find that the PRM can appear to adapt to a particular user while partly acting on population-level regularities rather than reliable evidence about that user's temporal state, helping some users and harming others without reliably identifying who has actually changed. Even the same held-out stable users receive much shorter assigned memory when the PRM is trained in a population where changes are common. Removing history length from the controller greatly weakens this behavior. Our findings show that *detecting change is not knowing whom to adapt, knowing whom to adapt is not knowing how to adapt, and adapting differently is not proof of genuine personalization*.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.