Same Causal System, Different Graphs: Auditing Pretrained Causal Discovery Across Measurement Coordinates
Abstract
Pretrained causal discovery systems predict graph structure for new datasets, but aggregate accuracy does not establish whether their reported relationships persist across measurement representations. We introduce OrbitBench, a paired evaluation protocol that re-expresses the same observations through strictly increasing, invertible transformations of each variable while keeping the inference pipeline fixed. We evaluate three pretrained systems, with PC and GES as classical references, on linear and nonlinear synthetic tasks and measurements from one physical light-tunnel apparatus. On 16 controlled linear tasks, non-decreasing skeleton F1 can coexist with changed adjacencies; among the pretrained systems, Arrow remains near-best in both views while changing its skeleton in 63 of 64 task–view comparisons. Disagreement persists when predicted edge counts are fixed and extends to directed-path answers on fresh tasks. On 32 new nonlinear tasks, using two audit views, audit-assisted and stability-first selection reduce macroaverage held-out adjacency disappearance risk from 8.27% for score-only selection to 6.71% and 6.68%, respectively. These gains concern persistence within the tested transformation families, not established improvements in causal correctness. Our results motivate reporting relationship consistency alongside accuracy and model-choice summaries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.