Proximal Causal Concept Explanations with Unobserved Concepts
Abstract
Concept-based explanations can reproduce a model's predictions yet yield biased causal attributions when unobserved concepts confound the target concept and model output. We introduce Proximal Causal Concept Explanations (PACE), which adapts proximal causal learning to estimate conditional mean concept effects without reconstructing unobserved concepts. PACE assigns two proxy measurements distinct roles: one enters an outcome bridge, while the other supplies conditional moments for learning it from factual observations. Under explicit causal, proxy, and bridge assumptions, averaged bridge contrasts identify these mean effects. We develop finite-feature moment estimators with optional predictive stabilization and characterize, for fixed features, how approximation, moment conditioning, and regularization affect population effect contrasts. Structural simulations demonstrate correction of confounding bias that predictive adjustment can retain. Analytical factor examples show that regression can recover a particular reference mean without recovering the bridge, while exact bridges can retain noise in individual contrasts. On CEBaB human-edited reviews, we evaluate intervention predictions for review-rating models using lexical features and pretrained language-model representations, including large language model embeddings. These same-review views have unverified proxy validity, and the comparisons show no consistent advantage over feature-matched regression. Together, these results distinguish bridge recovery, mean-effect accuracy, and individual-intervention prediction, highlighting the importance of measurement design and evaluation targets in causal concept explanation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.