acceptodds
Under review as a conference paper at ICLR 2027

Pseudo-Label-Aware Unlearning of Labeled Data in Self-Training

Abstract

Self training fits a labeling model on a small labeled dataset and uses its predictions to assign pseudo labels and select unlabeled examples for downstream training. After deletion of a labeled record, its influence may remain in this generated supervision. Removing the record from downstream training while keeping stored pseudo labels and selection decisions can therefore fail to reproduce full retraining without that record. We propose Pseudo label Aware Labeled data Machine Unlearning (PALMU). PALMU updates the labeling model on retained labeled data and regenerates pseudo labels for the unlabeled pool. Examples whose predicted class probability meets the confidence threshold enter downstream training. PALMU then updates the downstream classifier without fitting both classifiers to convergence. For fixed representations, we bound PALMU’s distance from full retraining using remaining downstream optimization error and supervision differences from the fully retrained labeling model. Across four datasets at deletion ratios of 5%, 10%, 20%, and 30%, PALMU meets predefined agreement and utility criteria in 397 of 400 fixed representation requests. Its mean predictive total variation is 4.3 to 56.6 times lower than the closest baseline evaluated against the same retraining reference. On CommitBERT, it retains 1.56 to 1.83 macro F1 points over labeled only retraining. For MLPs with learned representations, updating only the final classifier does not match retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.