Causal Structure Refinement with Warm-Start Reinforcement Learning and Finite-Sample Guarantees
Abstract
Causal structure learning from observational data is challenging due to the combinatorial search over directed acyclic graphs and the difficulty of reliable model selection under finite samples. Reinforcement learning (RL) offers a flexible framework for exploring the space of graph structures, but remains largely heuristic and lacks statistical guarantees. We propose a warm-start RL framework that recasts causal structure learning as refinement of an existing graph. Starting from an initial graph, the method performs a sequence of local edge modifications guided by a learned policy, while maintaining a dynamically constructed candidate set for selection. We adopt a copula-based BIC score that extends BIC to data with non-Gaussian marginals through a rank-based Gaussian-copula construction. We establish two results. First, under a monotone progress-potential condition, we bound the expected number of proposals required to reach a 1-optimal graph, with the bound scaling linearly in the warm-start depth. Second, the population-best fitted candidate in the explored archive is selected, using independent data, with high probability under a variance-adaptive Bernstein bound. Empirically, refinement improves the composite score of its initializer on all six benchmarks without increasing SHD, including the 100- and 223-node benchmarks where most baselines do not converge. Compared to RL without warm start, our approach achieves over improvement in Score, up to gains in TPR, and up to 50% reduction in structural error (SHD) on the two networks where this ablation was run.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.