acceptodds
Under review as a conference paper at ICLR 2027

Disentangled Counterfactual Diffusion for Identity-Preserving Explanations of Black-Box Medical Image Classifiers

Abstract

Complex medical image classifiers, particularly deep learning models, can achieve high predictive accuracy, yet often provide little evidence about which image factors truly drive their decisions, creating substantial safety risks. Saliency maps often highlight regions associated with a prediction but cannot test whether changing those factors would change the prediction; unconstrained generative counterfactuals may also alter patient-specific anatomy and thereby confound the explanation. We introduce Disentangled Counterfactual Diffusion, an identity-preserving framework that explains black-box medical image classifiers through controlled interventions designed to change their predictions. The model factorizes each image into a compact class-associated code and a multiscale individual representation. To construct a counterfactual, it retains the source individual representation, replaces the class-associated code with the target-class code, and performs global DDIM editing. Training combines diffusion reconstruction with class supervision, prototype alignment and separation, adversarial suppression of class information in the individual branch, and cross-covariance decorrelation. For counterfactuals generated through code exchange, additional objectives enforce the classifier prediction, class-code transfer consistency, preservation of individual features, anatomy, and edges, cycle reconstruction, and image realism. By comparing black-box classifier outputs and image differences between an original image and its identity-matched counterfactual, the method identifies class-associated factors that are sufficient to alter the model decision while controlling for subject-specific structure. We evaluate the method across multiple medical image classification tasks. The results show that, compared with existing approaches, our method more accurately identifies the key factors that influence model decisions and provides more precise explanations, offering an actionable route to understanding the decision process of black-box medical image classifiers.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.