acceptodds
Under review as a conference paper at ICLR 2027

HorizonBench: Long-Horizon Personalization with Evolving Preferences

Abstract

User preferences change over months of interaction, and a personalized assistant must recognize when a later life event supersedes a preference the user stated earlier. Natural conversations reveal neither the user’s current preferences nor the cause of a change, so this ability cannot be measured on real data. We introduce HorizonBench and its data generator, the first resource that explicitly links a user’s latent preference state to months of conversation and records the life event behind every preference change. The generator produces conversations from sampled state trajectories, tracks when each preference changes separately from when it is expressed, and calibrates how likely a life event is to change preferences, and by how much, against a 13-week longitudinal study with 160 participants. HorizonBench contains 4,245 questions from 360 simulated users, each asking for the response that best fits the user’s current preferences; when a preference has changed, one option reflects its earlier value. All 25 frontier models we evaluate choose this outdated option more often than expected if errors were spread evenly (29.2% of wrong answers against 25%), and the best model reaches 52.8% accuracy. The pattern holds across context lengths and expression styles, including when the preference was expressed shortly before the question, and retrieval and memory methods do not recover the lost accuracy, because the evidence of a change is a life event that never states the new preference. Remembering what a user said and knowing what the user prefers now are therefore separate capabilities, and current models lack the second. HorizonBench supports research on long-context models, memory systems, theory-of-mind reasoning, and user modeling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.