acceptodds
Under review as a conference paper at ICLR 2027

From Reasons to Revisions: A Tractable Theory of Belief Change for Boolean Classifiers

Abstract

Explanations of learned classifiers are typically one-directional: a model justifies a prediction, but disagreement with that explanation does not provide a principled way to revise the model. We propose a human-in-the-loop framework in which explanations can be contested and used to correct classifier behavior. Grounded in sufficient reasons for Boolean classifiers, correction is formulated as a minimal adjustment of the reasons supporting a prediction, connecting explainable machine learning with belief change. The main challenge is computational, since a prediction may admit exponentially many sufficient reasons, making explicit enumeration impractical. We show how, for a classifier compiled into an SDD, the compiled structure can be exploited to enumerate reasons directly according to the degree of change they induce. The procedure first identifies the minimum disagreement level and recovers all sufficient reasons obtaining it, while pruning candidates that exceed the current level; less conservative alternatives can then be generated on demand by progressively relaxing this bound. This avoids materialising the full set of sufficient reasons before identifying the most conservative corrections. Experiments on random and structured worst-case classifiers demonstrate the practical advantage of obtaining minimal corrections without first performing exhaustive prime-implicant enumeration. Since learned models such as decision trees can be compiled into these representations, our approach provides a symbolic correction layer for learned predictors and supports an interactive neuro-symbolic explanation-revision loop

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.