Personalization Done Right: Balancing Personalization and Privacy
Abstract
Personalized language models typically condition on user-specific context to better serve individual needs. However, the very data that enables effective personalization, such as user profiles, interaction histories, and preference signals, simultaneously increases privacy risk. Despite the growing importance of personalization and privacy in language modeling, their joint analysis remains limited. In this work, we study how personalization granularity affects task utility and user identifiability. We formalize a setting in which a user issues a prompt and the model has access to a user profile of varying granularity, and we characterize identity leakage through posterior identification gain. We then show that idealized group averaging reduces identity leakage under stated conditions, with utility change bounded by within-group utility differences. We study this tradeoff across a spectrum of personalization techniques, including retrieval-augmented generation, persona-based prompting, profile-augmented prompting, and provide extensive empirical findings that support its practical implications. Our findings have significant implications for the development of personalized language models that respect user privacy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.