acceptodds
Under review as a conference paper at ICLR 2027

From Membership Scores to Unlearning Claims in Offline Reinforcement Learning

Abstract

Machine unlearning aims to remove the influence of selected training data by matching a model retrained without them. Membership scores offer a convenient evaluation tool, but interpreting their changes requires knowing which properties of retraining they measure. We study this connection through controlled experiments in offline reinforcement learning, where simulator access and training records allow us to separate policy compatibility from training exposure. Naturally generated non-members receive member-like scores under a fixed audit, while interventions on reference selection and target adaptation reveal how additional information changes the conclusions of the same score family. We then examine whether score-based comparisons can recognize retain-only retraining itself. Two candidate audit designs frequently reject retraining controls despite consistently rejecting full-data comparators, showing why comparator separation alone does not validate an equivalence test. We explain these findings by distinguishing model-distribution distance from distances between individual models and by characterizing information lost in finite action-error scores. A three-action decision rule connects valid uncertainty bounds to support, rejection, or abstention, and a practical protocol specifies the claim, access assumptions, independent samples, and controls needed to apply it. The resulting case study shows how membership evidence can remain useful for diagnosis while supporting retraining-equivalence claims only through an explicitly validated statistical connection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.