Certified Feature Unlearning under Selective Feature Revocation
Abstract
Feature unlearning seeks to remove a model's predictive dependence on specified input features while preserving the utility of the remaining features. Existing feature-unlearning methods typically re-access the original values of the very feature being removed, which is often impermissible once that feature has been revoked for legal, privacy, or governance reasons. We study this selective feature revocation setting, in which the retained features and labels remain available but the per-example values of the revoked feature do not, and propose Bounded Mechanism-Covered Replacement (BMCR). BMCR replaces the revoked feature with values drawn from a declared bounded range and fine-tunes the model so that its predictions stay stable across these replacements; the declared range is the only information BMCR uses about the revoked feature, and its coverage defines the scope of the guarantee. We prove that the BMCR objective penalizes residual direct dependence on the revoked coordinate and that its near-minimizers are prediction-level close to retraining without that feature. We further show that, at any finite training checkpoint, the effect of off-manifold replacements on in-distribution risk and the model-assisted attribute-inference advantage over the retained features alone are both controlled by the same residual dependence. Experiments on seven tabular datasets, three architectures, and single-, multi-, and sequential-deletion settings show that BMCR matches retraining in utility, outperforms replacement and structural-masking baselines, and is about faster than retraining in wall-clock time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.