Model-Agnostic Invariant Causal Prediction Testing in Learned Representations
Abstract
Identifying invariant causal features in learned representations by modern machine learning models is central to robust generalization across environments and to the development of trustworthy AI systems. Based on invariant prediction, we develop a model-agnostic test of whether the conditional distribution of a response given a candidate feature subset remains invariant across environments. Under conditional invariance, the adjusted predictiveness scores coincide, even when the working prediction class is misspecified. We establish asymptotic normality of the component statistics and a chi-square null limit for their squared cyclic aggregate, providing asymptotic Type I error control. Power depends on whether the chosen working class and score distinguish changes in the conditional mechanism. Our approach can be readily integrated with pretrained foundation models. Simulation studies show that the proposed test reliably identifies causal features in complex settings and remains robust to severe model misspecification. Experiments using the Contrastive Language-Image Pretraining (CLIP) model further illustrate that our approach uncovers meaningful invariant features and improves generalization in downstream tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.