acceptodds
Under review as a conference paper at ICLR 2027

Unlearning Offline Stochastic Multi-Armed Bandits

Abstract

Machine unlearning aims to unlearn data points from a learned model, offering a principled way to mitigate privacy risks without full retraining, while preserving the utility of the learned model. Prior work has mainly studied unsupervised / supervised machine unlearning, leaving unlearning for sequential decision-making systems far less understood. We initiate the first study of a foundational sequential decision-making problem: offline stochastic multi-armed bandits (MAB). We formalize the privacy constraint for offline MAB and measure utility by the post-unlearning decision quality. We conduct a systematic study under two data-generation models: the fixed-sample model and the distribution model. Our analysis covers single-source requests in both models and extends to multi-source deletion in the fixed-sample setting. For these settings, our algorithmic design is built on two base mechanisms: the downward-shifted Gaussian mechanism and rollback, and we propose adaptive algorithms that switch between them according to the data regime and privacy constraint. We also provide performance guarantees across the above settings and establish lower bounds under both dataset models. Experiments validate the predicted trade-offs and show the effectiveness of the proposed methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.