A Personalized Language Model Should Be Less Surprised by You
Abstract
Personalization in language models is commonly evaluated through task performance, stylistic resemblance, or automatic judgments. These proxies capture useful behaviors, but they do not test whether user-specific conditioning helps a model better anticipate its target user, as reflected in lower surprisal on held-out user text. For a fixed task, tokenizer, and base model, we define predictive personalization as the extent to which a user-specific adaptation reduces the surprisal of text produced by the target user relative to a generic baseline. We operationalize this notion as user-specific predictive alignment and introduce Likelihood Improvement from Targeted Familiarity (LIFT), a per-token log-likelihood-ratio diagnostic that measures the length-normalized surprisal reduction on held-out user text relative to a generic baseline. We also introduce ESS-LIFT as a rare-token robustness analysis that downweights rare, high-leverage tokens to test whether conclusions drawn from LIFT depend on their influence. Since there is no scalar oracle for the latent amount of personalization achieved by an adaptation, we evaluate LIFT through construct-validity stress tests, demonstrating its sensitivity to user identity beyond contextual relevance, robustness to generic fluency and raw-likelihood confounds, and ranking coherence across qualitatively different personalization mechanisms. Experiments across diverse language models, datasets, and tasks provide convergent evidence that LIFT captures user-specific predictive alignment, supporting its use as a principled diagnostic for predictive personalization in language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.