Provable Guarantees for Learning Pure Latent High-dimensional Structural Equation Models
Abstract
Learning Directed Acyclic Graphs (DAGs) from observational data is a fundamental challenge in causal inference, exacerbated when the true causal mechanisms occur among unobserved latent variables. In this paper, we study the identifiability and learning of pure latent high-dimensional Structural Equation Models (SEMs). While traditional causal discovery algorithms only recovering Markov equivalence classes (CPDAGs) or succumbing to latent confounding. We propose a generic solver framework to recover the exact, fully-directed latent DAG in pure latent high-dimensional SEMs. Our two-phase approach first reconstructs the latent covariance from observational data, and then iteratively identifying and marginalizing terminal vertices to recover the latent DAG structure. Under this framework, we relaxes the faithfulness assumption in standard causal discovery algorithm, and provides theoretical guarantees for topological ordering, bounded parameter estimation error, and exact support recovery via thresholding. We establishing that exact latent DAG recovery is achievable with logarithmic sample complexity of where is the number of latent variable. Our experiments on synthetic data and a real-world fMRI brain connectivity dataset demonstrates that our method successfully recovers exact directed causal structures where standard observed or equivalence-class baselines fail.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.