acceptodds
Under review as a conference paper at ICLR 2027

Steering Circuit-Level Activations for Personalized Knowledge Injection

Abstract

Although large language models (LLMs) are increasingly personalized to individual users, they still struggle with a deceptively simple but critical requirement: updating mutable personal facts, such as a user's job, location, or relationships. Fine-tuning approaches are costly and prone to catastrophic forgetting. Meanwhile, knowledge editing approaches rely on subject-bound representations; they fail in personal domains because private entities lacking prior exposure lead to unstable key extraction, while multi-attribute updates cause key collisions. We propose SPIKE (Steering Circuit-level Activations of Language Models for Personalized Knowledge Injection), a methodology that identifies the LLM components effective for injecting personal knowledge and incorporates mutable personal knowledge into the LLM through activations generated via a KG-LLM Alignment Module. Our experiments demonstrate that SPIKE effectively balances the accurate incorporation of new facts with the preservation of existing knowledge, offering a practical solution for personalization in settings where user information evolves over time. Our code is available at https://anonymous.4open.science/r/spike.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.