CausalMix: Defining and Benchmarking Hybrid Causal Discovery
Abstract
Classical causal discovery uses numerical observations and interventions, while large language models (LLMs) also bring knowledge associated with variable names. A graph-recovery score alone does not reveal which information source supported the prediction. We introduce **CausalMix**, a simulation and evaluation suite for hybrid causal discovery that compares names-only, data-present, anonymized, and classical data-only conditions under controlled graph and data budgets. The benchmark separates output validity, name effects, local responses to data, and accurate recovery. Our evaluation spans six reference networks and four synthetic graph configurations with 5–50 variables. On the evaluated reference graphs, recovery is strongly name-mediated, while gains from numerical evidence are small or unstable. These name effects may reflect domain knowledge or benchmark recall; our controls do not distinguish the two. We also post-train Qwen3-4B with supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO). Post-training improves output validity, but adding data improves Sachs recovery only when names are anonymized, and recovery on 20–50-variable graphs remains far below a classical interventional baseline. CausalMix provides a reproducible protocol for testing whether hybrid methods use supplied evidence reliably, rather than inferring evidence use from aggregate scores.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.