A Paired Case Study of Conditional and Marginal Feature-Replacement Scores
Abstract
Feature-removal explanations measure how a predictor responds when an input feature is replaced. Because the predictor still requires a complete input, every such explanation depends on a replacement distribution. Conditional replacement is commonly motivated by its ability to preserve dependencies among observed features, but observational fidelity does not by itself determine how the predictor will respond to the resulting inputs. We study this gap with a paired evaluation that holds the predictor, evaluation examples, labels, and scoring rule fixed while changing the replacement distribution. The primary study compares conditional and marginal replacement on Dry Bean and evaluates coordinate prediction, covariance preservation, score magnitude, and feature ordering. Complementary studies examine stochastic draws, deterministic means, and label-informed replacements across several tabular domains. Conditional replacement more closely reproduces observed coordinate and covariance structure, yet it produces substantially different feature-removal scores and rankings from marginal replacement through the same predictors. The secondary comparisons further show that sensitivity to the replacement procedure depends on the dataset and predictor family. These findings establish the replacement distribution as part of the feature-removal estimand rather than a secondary implementation detail. A reproducible explanation should therefore report the replacement rule together with the predictor and scoring procedure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.