Scaling Temporal Causal Discovery to Millions of Variables
Abstract
Temporal causal discovery becomes prohibitively expensive as the number of variables grows: with N series and L lags, a method may need to consider N^2L candidate relationships. Yet many real temporal systems are sparse, with each variable depending directly on only a small fraction of the system. We exploit this structure with VL-PCMCI, a scalable causal discovery framework that separates cheap global screening from expensive causal reasoning. VL-PCMCI first screens all lagged pairs using GPU-accelerated matrix operations, performs causal discovery independently within sparse blocks, and then reconciles relationships across blocks. We characterize when this decomposition preserves causal decisions and identify three requirements: true causes must survive screening, tested non-edges must be given separating sets, and statistical testing must account for screening. Experiments show that VL-PCMCI maintains high accuracy while scaling substantially beyond existing temporal causal discovery methods, recovering nearly every true edge in planted systems with millions of variables. At this scale, computation is no longer the only bottleneck: false positives from data-dependent screening grow approximately quadratically unless the final tests account for selection. These results show that temporal causal discovery at millions of variables is computationally feasible while identifying the statistical conditions required for reliable discovery at this scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.