KEY: Unlocking Disentangled Model Updating with a Knowledge Addressing Dictionary
Abstract
Recent years have witnessed growing interest in knowledge updating (e.g., editing and unlearning) in language models. Despite substantial progress, existing methods remain challenged by knowledge entanglement, where different knowledge items share overlapping parameters or representational directions, often causing unintended interference. To address this issue, this paper introduces KEY (Knowledge addrEssing dictionarY), aiming for unlocking distinguishable knowledge structures within the entangled weight space of language models. Specifically, the proposed method first constructs an orthogonal atom dictionary from knowledge-induced gradients by decomposing their input- and output-side directions. It then performs sparse addressing to identify a compact set of atoms and their coefficients for each update. Theoretically, both global and local proximity bounds are established, characterizing conditions under which the learned dictionary preserves the desired knowledge structure and supports stable knowledge addressing. Extensive experiments across both knowledge editing and unlearning tasks demonstrate that KEY consistently achieves competitive performance against recent state-of-the-art baselines, while enabling more interpretable and disentangled knowledge updating.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.