Data-Aware Sensitivity Quantification of Tree Ensembles
Abstract
Decision tree ensembles (DTEs) are widely used for classification in domains such as finance and healthcare, motivating the need to formally analyze their behavior. One important property is sensitivity, where changes to a designated subset of input features can alter the model’s prediction. Most existing work studies quali- tative notions of sensitivity, while the few quantitative approaches measure sensi- tivity over the input space without accounting for how likely different inputs are under the underlying data distribution. In this work, we introduce a data-aware notion of quantitative sensitivity. Our input consists of a DTE together with a Bayesian network capturing the distribution of the underlying data. We first iden- tify the regions of the input space that exhibit sensitivity and then weight each region by its probability, so that the measure reflects how likely sensitivity is to occur in practice, and changes with the data distribution even though the model does not. We develop a knowledge-compilation-based algorithm that computes weighted sensitivity within certified error and confidence bounds, and implement it in our tool Weighted-XCount, which scales well beyond direct weighted model counting and shows that weighted and unweighted sensitivity can differ sharply.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.