PersJEPA: Interpretable Language Model Personalization through Sparse Residual Prediction
Abstract
Language models can personalize their answers when given a user's profile or past interactions, but repeatedly supplying this information consumes context and computation. Learning personalized parameters can reduce this dependence on prompting, yet offers limited visibility into how user preferences affect the model's responses. Activation steering has attracted growing interest as a way to personalize models through interventions in hidden states without updating their parameters, yet it faces three key challenges: (1) fixed adjustments overlook how user preferences interact with each request, (2) steering confined to the initial decoding step leaves later steps without further intervention, and (3) accurate prediction of hidden representation changes does not reliably translate into personalized responses. To address these challenges, we introduce PersJEPA, a framework for interpretable personalization of frozen language models through sparse prediction conditioned on user information. Inspired by JEPA, we formulate predictive residual learning, using paired generic and personalized representations to learn how user context changes the model's hidden states. We design sparse residual decoding so that each request activates a small set of learned features whose directions are combined into a personalization residual. User conditioning determines which features contribute and how they adjust the model's representations. We further develop generation-aligned steering, applying the decoded residual throughout generation and optimizing the predictor for the likelihood of personalized responses through the frozen model while retaining residual supervision. The language model remains frozen throughout training, and inference requires no profile text in the prompt. Experiments demonstrate substantial improvements in controlled persona generation across multiple language models and gains on real user personalization benchmarks. Analyses of the learned features reveal how their contributions vary across user groups, providing a concrete basis for interpreting the learned personalization interventions. Our implementation is fully open-sourced.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.