Vocabulary Aligned Explanation for Computer Vision Models
Abstract
Optimizing vision models for accuracy does not force the model to learn and utilize human-understandable and domain-specific concepts and rules. We propose explaining image classifications through domain-specific concepts, visual evidence, and expert knowledge expressed as logical rules. We introduce FLUENT (Framework for Localized, Uncertainty-Aware Explanations in Normative Terminology), a post-hoc framework that ranks expert-authored rules combining jointly required, alternative, and absent concepts. To account for prediction errors, FLUENT combines predicted labels and concept states with empirically estimated likelihoods and concept priors. It scores each rule by calculating support over satisfying assignments to the concepts mentioned in that rule, then normalizes these scores to express relative confidence among candidate explanations. The selected rule is illustrated through occlusion-based highlights for positive concept mentions in the query image and localized examples in retrieved reference images for negative mentions. These illustrations are intended to show end users unfamiliar with the domain what the concepts named in an explanation look like. FLUENT combines domain knowledge and model accuracy, and delivers visual concept highlights and explicit rule confidence to explain an image in the vocabulary of an application domain.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.