URGE: A Unified Representation-Guided Loss Reweighting Framework for Fine-grained Machine Unlearning
Abstract
Machine unlearning is an emerging paradigm to remove specified data influences from trained models, with broad applications across image classification, diffusion models and large language models. A prominent direction is gradient-based unlearning (GU), which designs forgetting objectives and retain regularization to balance forgetting and utility. Loss reweighting further improves GU by explicitly assigning different update strengths to samples or tokens based on their forgetting needs. However, existing methods face two key limitations: (i) effectiveness: prediction-distribution-based methods rely on signals that are partially redundant with the original optimization gradients; and (ii) applicability: other methods require external auxiliary models, retain data, or task-specific designs. In response, we propose **URGE**, a **U**nified **R**epresentation-**G**uided r**E**weighting framework for GU. URGE constructs weights for each sample/token from two representation-based signals: Conditional Representation Invariance estimates its intrinsic need for forgetting, while Representation Similarity estimates its remaining need for further forgetting. By leveraging internal representations, URGE captures richer information about how samples/tokens are encoded in the model, improving unlearning effectiveness. Without requiring external models or retain data, URGE is applicable across GU methods, tasks, forget targets, and granularities. Experiments against 18 baselines across 3 tasks and 13 model-dataset settings show that URGE generally achieves a better forgetting-utility trade-off than GU and reweighting baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.