RECAP: Relative Evaluation with Consensus Augmented Policy Search for Causal Discovery
Abstract
Reinforcement learning can search for causal graphs by scoring sampled structures, but learning a graph policy from limited candidate graphs is often unstable. Rewards vary across inputs, while promising edges found in earlier updates may be lost as the policy changes. We propose RECAP, Relative Evaluation with Consensus Augmented Policy Search, which couples within-input relative evaluation with cross-update structural memory. RECAP uses group-relative advantages to distinguish candidate graphs for the same input and guide the current policy update. To carry structural evidence across updates, it aggregates edges from high-reward graphs into an evolving consensus teacher, whose selective, annealed guidance favors recurring edges without permanently constraining the search. A reference-policy KL term moderates policy changes, while the graph policy samples candidates from contextual variable representations and scores them using a BIC-based reward with an acyclicity penalty. Across four benchmarks spanning Bayesian networks, simulated fMRI, and protein signaling, RECAP achieves the lowest SHD and highest precision, with the best or tied-best F1 on three. Ablations confirm the contributions of both components, while equal-budget comparisons further demonstrate RECAP's F1-SHD advantage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.