A Token Rendering Perspective on Large Language Model Unlearning
Abstract
LLM unlearning seeks to remove targeted knowledge without disrupting retained capabilities. We argue that many existing output-based methods control which predictions to suppress without explicitly parameterizing the contextual relations involved in producing them. We therefore introduce a token rendering perspective, which views autoregressive generation as a hierarchical rendering process in which contextual information is transmitted through attention-mediated token relations and progressively rendered into next-token probabilities. This perspective provides a common first-order characterization of several representative unlearning objectives as forms of target-token renderability control, highlighting that these objectives do not explicitly parameterize the underlying token relations. Based on this view, we propose -Unlearning, which freezes the base LLM and explicitly tunes the renderability of token-token relations. We estimate the sensitivity of semantic relations to forget and retain rendering, attenuate forget-sensitive relations, and constrain changes to retain-sensitive ones. For efficient intervention, relation transmittance is parameterized over semantic token clusters. Experiments on TOFU and WMDP across multiple LLMs and forgetting ratios demonstrate that -Unlearning achieves competitive forgetting–utility trade-offs under the evaluated settings. Code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.