GRACE: Population-Level Adversarially Robust Counterfactual Explanations
Abstract
Counterfactual explanations (CFs) suggest changes to an input to flip a machine learning model's prediction toward a desired outcome. Most existing CF generators are not robust because they assume a static model and data distribution. However, models deployed in practice are retrained periodically, so a CF that is valid today may not be so tomorrow. Institutions typically issue a population of CFs, making robustness a population-level property, rather than an individual one. Existing robustness methods, however, optimize each CF against its own worst-case perturbation rather than the shared population-level risk induced by retraining, which may cost more proximity than necessary. We introduce Geometrically Restricted Adversarially robust Counterfactual Explanations (**GRACE**), a distributionally robust wrapper that treats the population of CFs produced by any CF generator as the robustness target. GRACE controls a surrogate for the population's worst-case invalidity under future retraining while keeping the refined CFs close to both the original CFs and the observed distribution. Across tabular and image datasets, GRACE substantially improves the future validity of the base CFs with only a small increase in proximity, matching or exceeding the future validity of recent state-of-the-art robust CF baselines on tabular benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.