acceptodds
Under review as a conference paper at ICLR 2027

Suppression Is Not Deletion: Evaluating Offline-RL Unlearning Against Retraining

Abstract

Deleting trajectories from an offline reinforcement learning dataset should leave the policy that training without them would have produced. Rather than retrain, several published offline-RL unlearning methods keep the deleted trajectories in the update and push the learner away from them, which we call suppression. This paper contributes a behavioral audit of unlearning against retraining, a per-state analysis of what suppression objectives optimize, and an evaluation of published methods on the benchmark released with TrajDeleter. Since retraining is random, a fresh seed giving a different policy, the audit compares a candidate with a family of independent retrains, and first audits the unchanged policy, the original with no unlearning applied, because if it passes, a candidate's pass proves nothing about deletion. At the benchmark's own deletion sizes the unchanged policy behaves like the retrained policies wherever the audit is validated on held-out retrains, so a pass there cannot show that a method deleted anything. On larger HalfCheetah deletions that do change behavior, no suppression method we evaluate reliably reaches the retrain family, whereas retrained policies stay inside it under further retained-data training. The analysis explains why: deletion gives the removed data zero weight, negative weighting of them overshoots the retraining target, and negating their rewards changes what the value function, the critic, learns; tests declared in advance find each effect in the predicted direction on HalfCheetah. The benchmark's own forgetting score follows the critic's values rather than the policy's behavior: shifting the critic by a constant reports complete forgetting while every action stays unchanged. Unlearning in offline RL should be judged against the behavior of retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.