Conditional Independence Testing from Dependent Samples with Applications in Causal Discovery from Time-series Data
Abstract
Causal discovery from time-series data is a fundamental problem with many applications in several industries. Most causal discovery algorithms use conditional independences, explicitly or implicitly, between observed variables to learn the causal graph. Directly applying these algorithms to time-series data is challenging due to the dependence across samples created through the causal effects over time. The landmark PCMCI algorithm addresses this issue by a two-stage approach: While the first stage is a conservative edge pruning stage, the critical second stage tests independence between two variables by conditioning on the union of parents of both variables. This removes the dependence between samples through time, allowing us to plug in conditional independence tests designed for IID data. However, large conditioning sets in the second stage often have low statistical power with finite data since few samples are left per stratum. To address this issue, we first analyze the chi-square independence test from dependent samples. For stationary, rapidly mixing processes, we show that the test statistic still converges to a chi-square distribution when conditioning on the parents of one variable, allowing the other variable to be dependent through time. Our work, to the best of our knowledge, provides the first explicit dependence of the chi-square test on the mixing time of the underlying time series. Based on this, we develop a causal discovery algorithm, which, in its second stage, conditions on the parents of only one of the variables. The smaller conditioning set significantly improves statistical power and causal discovery performance. We show through synthetic experiments that our algorithm outperforms PCMCI and other baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.