acceptodds
Under review as a conference paper at ICLR 2027

Counterfactual Fragility Certificates for Tabular Classifications: Guarantees and Limits

Abstract

A high-confidence tabular prediction need not survive unreliable input fields. We study protocol-relative fragility: loss of support for the original prediction under a declared finite collection of field failures. Redundant fields can mask one-step effects, while partial degradation can differ from deletion. Counterfactual Fragility Certificates (CFC) combine a sequential removal path, recomputed after every step, with graded tests of every group. An original-class margin records collapse even after the predicted class changes. A reconstructable record exposes the tests, decisions, and witnesses behind a fixed fragility dominance score (FDS). We distinguish record verification from worst-case robustness, causal recourse, and prediction correctness. A conditional greedy bound and counterexamples delineate generic search limits; a model-specific affine solver certifies minimum removal depth. We evaluate a released implementation on seven OpenML datasets, eight classifier families, and three seeds, with paired row-level outputs and reference checks. On the in-family endpoint, CFC-FDS achieves 0.7743 AUROC versus 0.7741 for prespecified, query-matched random search with the same three-term score; the paired difference is +0.0001 (95% interval [-0.0054, 0.0063]). On withheld donor/quantization mechanisms, the difference is -0.0040 ([-0.0089, 0.0004]). These comparisons do not establish a consistent advantage for adaptive removal over query-matched random search across both endpoints. A separate five-dataset restoration study measures 2.36 net errors averted per 100 cases at 10% review using a validation-trained utility ranker, versus 2.61 for raw confidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.