Learning Causal Variable Orderings with Pointer Networks for Causal Structure Discovery
Abstract
Causal structure learning from observational data requires searching over a large combinatorial space of possible directed acyclic graphs (DAGs), making accurate and efficient structure learning challenging. Existing causal discovery methods based on continuous optimization and reinforcement learning either optimize graph structures directly or explore variable orderings through sequential decision making. We propose CaPo, a neural causal discovery framework that formulates causal variable ordering as a learnable permutation to guide constrained graph search. A Pointer Network generates candidate variable orderings, while an ordering-induced mask restricts admissible edge directions according to the learned ordering, thereby reducing the search space over candidate graph structures. The key novelty is to learn variable orderings directly from observational data and optimize them according to downstream causal graph quality, rather than relying on fixed or explicitly searched orderings. The ordering policy is optimized through reinforcement learning using a graph-level causal structure reward, and a DAG-constrained graph learner estimates the corresponding causal structure under the learned ordering. Experiments on Bayesian network benchmarks of varying sizes show that CaPo achieves competitive or lower Structural Hamming Distance (SHD) and higher F1-scores than state-of-the-art causal discovery baselines across multiple datasets. CaPo also requires less computational time than several compared methods, demonstrating efficient causal structure learning across networks of different sizes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.