TSCU: Topology-Guided Subspace-Constrained Unlearning for Large Language Models
Abstract
Recent advances have strengthened the reasoning and knowledge capabilities of large language models (LLMs), but their capacity to retain information raises concerns about memorized private, copyrighted, or harmful content. Machine unlearning is therefore important for trustworthy LLM deployment. Existing approaches often suppress target outputs or directly modify model parameters. However, output suppression may leave internal recovery pathways intact, whereas non-localized parameter updates can impair retained knowledge and general capabilities. This motivates localized unlearning that identifies where target knowledge is represented and constrains interventions accordingly. We propose TSCU, a topology-guided, subspace-constrained framework for representation-level LLM unlearning. TSCU uses layer-wise topological discrepancies between forget and knowledge-attenuated reference representations, together with retain stability, to identify forgetting-sensitive yet retention-stable layers. It then estimates forgetting-relevant subspaces and restricts low-rank interventions to these directions. A dual-branch alignment objective moves forget representations toward knowledge-attenuated states while restoring perturbed retain representations. Experiments across multiple settings show that TSCU improves forgetting robustness and locality while preserving retained knowledge and general utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.