From Reasons to Revisions: A Tractable Theory of Belief Change for Boolean Classifiers
Abstract
Explanations of learned classifiers are typically one-directional: a model justifies a prediction, but disagreement with that explanation does not provide a principled way to revise the model. We propose a human-in-the-loop framework in which explanations can be contested and used to correct classifier behavior. Grounded in sufficient reasons for Boolean classifiers, correction is formulated as a minimal adjustment of the reasons supporting a prediction, connecting explainable machine learning with belief change. The main challenge is computational, since a prediction may admit exponentially many sufficient reasons, making explicit enumeration impractical. We show how, for a classifier compiled into an SDD, the compiled structure can be exploited to enumerate reasons directly according to the degree of change they induce. The procedure first identifies the minimum disagreement level and recovers all sufficient reasons obtaining it, while pruning candidates that exceed the current level; less conservative alternatives can then be generated on demand by progressively relaxing this bound. This avoids materialising the full set of sufficient reasons before identifying the most conservative corrections. Experiments on random and structured worst-case classifiers demonstrate the practical advantage of obtaining minimal corrections without first performing exhaustive prime-implicant enumeration. Since learned models such as decision trees can be compiled into these representations, our approach provides a symbolic correction layer for learned predictors and supports an interactive neuro-symbolic explanation-revision loop
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.