acceptodds
Under review as a conference paper at ICLR 2027

Forget Without Seeing: Unlearning Behaviour Transfer without Access to the Forget Set

Abstract

Transferring unlearning ability to personalized models poses a practical dilemma: delivering the data to be forgotten exposes them anew, while directly replacing the downstream model with the upstream unlearned model or applying the upstream parameter update may disrupt locally acquired capabilities. We study a Forget Without Seeing problem, in which downstream models inherit upstream unlearning behaviour without receiving the original forget set during the unlearning update. We propose Carrier-mediated Unlearning Behaviour Transfer (CUT), inspired by subliminal learning. CUT encodes the behavioural change induced by upstream unlearning into a carrier of synthetic non-target text, enabling personalized model owners to learn endpoint-specific updates without accessing the forget set. CUT consists of three stages. First, a Teacher-Matched Steering Vector is learned to align the Original Model with the Unlearned Teacher on forget and retain calibration data. Second, the same Original Model generates paired responses to synthetic non-target prompts with and without steering, forming a Paired-Text Carrier. Finally, each downstream model learns an endpoint-specific update from the carrier, transferring unlearning behaviour while preserving personalization. Only the text carrier is delivered downstream; the forget/retain data, Unlearned Teacher, Teacher logits, and steering vector remain upstream. We evaluate unlearning transfer on TOFU across Qwen2.5-7B, Gemma-7B, and Llama-2-7B, and personalization preservation on LaMP-3. CUT transfers upstream unlearning behaviour across algorithms and forget-set sizes. On Qwen2.5-7B with SimNPO, it achieves unlearning comparable to directly applying the upstream parameter difference while reducing personalization mean absolute error (MAE) by approximately 20%. Privacy evaluations further show that the text carrier limits target-information exposure under the tested auditing, reconstruction, and update-role inference attacks. These results demonstrate that unlearning behaviour can be transferred without redistributing the original forget data, while preserving downstream personalization and limiting information exposure during delivery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.