acceptodds
Under review as a conference paper at ICLR 2027

CoForget: Composing Heterogeneous Unlearning Requests to Activate Dormant Backdoors

Abstract

Unlearning-activated backdoor attacks keep a backdoor dormant after training and activate it later through an unlearning request. Requests composed entirely of camouflage samples can induce this transition but may also reduce benign accuracy. Since clean samples can influence backdoor behavior, incorporating them alongside camouflage samples provides an additional way to improve this balance. To this end, we introduce CoForget, a request-composition attack that combines a small fixed camouflage component with deliberately selected clean samples. CoForget first calibrates a universal trigger with a logit gap objective to promote dormancy without excessively suppressing the target response. It then greedily selects a clean subset whose cumulative gradient aligns with an activation direction estimated on a local shadow model. Neither trigger nor request construction requires access to the victim model or knowledge of the provider's unlearning algorithm. Experiments across three image classification datasets and three approximate unlearning algorithms show low attack success before unlearning and high success afterward. In the evaluated clean component ablations, selected samples improve mean benign accuracy over camouflage-only requests and retain high activation. Comparisons at matched total request sizes further show strong activation while using substantially fewer camouflage samples than the evaluated baselines. These findings motivate considering request composition when assessing unlearning security.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.