Do Language Models Have a Model of You? Discovering and Enhancing Latent Representations of User Personas
Abstract
Humans carry working models of the people we know, and these models are useful because they persist across contexts and support predictions about future behavior. Language-model personalization is usually evaluated very differently: almost entirely by whether the model produces an adapted response. This creates a central blind spot. A system can achieve strong personalized output quality while maintaining a user state that is weakly recoverable, poorly calibrated, or unstable across different histories of the same person. We therefore distinguish behavioral personalization, which concerns whether outputs look adapted to a user, from representational personalization, which concerns whether the system maintains a recoverable and reusable user-state estimate that explains behavior across interactions. We formalize representational personalization through persona identifiability: a useful user state should be predictive of held-out behavior, recoverable from past interactions, and stable across partial histories of the same user. To study this distinction, we introduce \benchmarkname, a controlled diagnostic protocol built from HumanLM user profiles and interaction records under a strict public/private split, and \methodname, an EM-inspired latent refinement procedure that couples behavior prediction with persona inference over explicit user-state estimates. Across controlled recovery, calibration, history-stability, user-shuffling, long-history, and real-user personalization evaluations, we find that output quality and user-model quality are related but not equivalent: models with similar or strong downstream personalization can differ substantially in the recoverability and stability of their inferred user states. Moreover, explicitly optimizing for identifiable user states improves recovery, calibration, cross-history consistency, and held-out user behavior prediction while remaining competitive or stronger on downstream personalization tasks. These results suggest that personalization should be evaluated not only by what a model says, but also by the quality of the user model that supports its behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.