acceptodds
Under review as a conference paper at ICLR 2027

Causal Discovery from Diverse Data

Abstract

Causal discovery is increasingly applied across heterogeneous environments whose data-generating processes may differ substantially. Unlike standard multi-domain formulations based on hard or soft interventions, we do not assume that the domains arise from interventions on a common baseline graph: their causal graphs may differ structurally, including through edge additions, deletions, or reversals. We ask a fundamental question: how much causal knowledge can be learned by jointly exploiting such heterogeneous datasets? We show that mechanisms that remain stable across domains can induce cross-domain distributional invariances, even when other mechanisms and the underlying causal graphs differ. By systematically leveraging these invariances together with within-domain conditional independences, we identify the distributional constraints that such invariances impose on tuples of domain-specific causal graphs and characterize the resulting inv-Markov equivalence class. In particular, cross-domain invariances can reveal causal orientations that are not identifiable from any domain in isolation. Our algorithm can leverage cross-domain invariances to learn cause-effect relations beyond what is learnable in each domain, even when the underlying causal graphs are arbitrarily different, such as having reversed edges. Building on this characterization, we develop Diverse-FCI, a constraint-based algorithm that learns domain-specific partially oriented graphs while exploiting pairwise cross-domain invariances. Under a faithfulness condition ensuring that the relevant cross-domain dependencies remain detectable after latent projection, we prove its soundness. Synthetic experiments show improved orientation recovery over learning each domain independently, with larger gains as structural heterogeneity increases.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.