Probing and Steering Vision Model Decisions along Concept Axes in Diffusion Space
Abstract
We introduce Mask-supervised Orthonormal Concept Axes (MOCA), a diffusion-based framework for concept-wise probing and steering of vision model decisions. Through independent manipulation of individual visual concepts in a given image and observing the resulting model response, MOCA enables concept-wise probing underlying a vision model's decisions. By directly scaling concept strengths, MOCA formulates visual explanation as concept-wise interventions, probing feature sensitivity to reveal whether specific visual elements explain the decision or act as spurious correlations. Experiments across multiple datasets and vision models demonstrate that MOCA produces spatially localized and semantically meaningful interventions, enables identification of spurious cues, and supports downstream mitigation. We extend MOCA for visual counterfactual explanation, providing a broader range of explanations for the same prediction through interventions on different visual concepts within the image. Experiment results show that MOCA improves the overall accuracy in the mitigation task by 44.9% on HardImageNet and 3.0% on COCO relative to the corresponding baselines, achieves comparable performance in counterfactual explanation tasks, and consistently support additional explanations through non-ground-truth concepts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.