acceptodds
Under review as a conference paper at ICLR 2027

Rotation-First Unlearning: Selective Weight Damage in a Data-Dependent Basis

Abstract

After training, a language model may need to forget specific content without losing the capabilities learned from the rest of its data. Most LLM unlearning methods try to achieve this through additional optimization, using objectives that suppress the forget set while regularizing behavior on retain data. We propose Rotation-First Unlearning (RFU), a training-free weight update built on a simple geometric premise: if forget and retain activations concentrate energy in different subspaces, the forget-heavy subspaces can be identified and attenuated directly. RFU runs one forward pass over forget and retain data, forms activation moment matrices at each target layer, and solves a generalized eigenproblem that orders directions by forget-over-retain activation energy. It then converts this ordering into a Euclidean-orthonormal rotation that preserves the leading generalized eigenspaces, shrinks the corresponding weight coordinates with a closed-form rule, and maps the weights back to their original parameterization. The RFU update is deterministic, label-free, and gradient-free. We analyze RFU with a first-order bound on the retain damage, and derive a spectral concentration diagnostic that predicts when the RFU update would be selective before any evaluation. Across TOFU, MUSE, and WMDP on Llama-, Qwen-, Gemma-, and Mistral-family models, standalone RFU improves forget-retain tradeoff in spectrally separable settings, and consistently enables deeper forgetting at matched or near-matched retention when used as a warm start for other training-based unlearning methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.