acceptodds
Under review as a conference paper at ICLR 2027

Certified Counterfactuals for Neural Networks

Abstract

Counterfactual explanations aim to identify minimal changes that alter a model’s prediction. Unlike adversarial perturbations, however, these changes should correspond to plausible, interpretable, and actionable modifications of the input. Existing approaches broadly fall into two categories: data-based and model-based methods. Data-based methods search among reference points, favoring plausibility but potentially sacrificing proximity and disclosing individual reference points. Model-based methods require query-time model access and generally provide no formal guarantee that the target prediction persists throughout a neighborhood of the returned point or that the point lies in a well-supported region of the input space. We introduce CertCF, a counterfactual explanation method for neural network classifiers that bridges these two families by linking counterfactual generation with neural network verification. CertCF uses a reference dataset to construct a certified under-approximation of the target-class preimages as a union of convex regions, and uses this structure at query time to enable fast, plausible, and valid-by-design counterfactual search. Additionally, it can guarantee bounded robustness to input perturbations and restrict changes to a designated subset of features. We evaluate CertCF on seven tabular datasets commonly used in the counterfactual explanations literature, comparing it against four representative methods, and show that CertCF achieves 100% validity and strong robustness while remaining competitive on the main quality metrics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.