Bounded-State Continual Personalization of LLMs under User Preference Drift
Abstract
Continual personalization is often studied under idealized assumptions with explicit preference signals or offline task boundaries. In realistic deployment, however, systems must learn from streaming user interactions, where preferences are implicit and change points are never annotated. A central difficulty is deciding when recent adaptation should alter the model used for subsequent interactions: predictive mis- match may reflect a persistent preference shift, a temporary deviation, or a difficult input. We introduce Stable-and-Fast (SAF), a bounded state personalization method that separates continuous adaptation from selective consolidation. A Stable adapter serves predictions, while a Fast adapter learns from each revealed interaction. Be- fore updating on the current feedback, SAF compares their predictive losses and accumulates relative evidence to consolidate Fast into Stable or reset Fast. We formulate this evidence as a length-normalized predictive log ratio and characterize its interpretation and limits under an evolving Fast adapter. We evaluate under a deployment-like streaming protocol on four LaMP derived tasks with 100 users per task, where SAF achieves good result compared to SOTA methods. These results establish selective consolidation as a distinct control decision for continual personalization in realistic streaming deployments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.