Where the Accuracy Goes in Single-Edit Causal Search
Abstract
Score-based structure learning returns the graph a greedy search reaches by repeatedly taking the single edit that most improves a score such as BIC, so its output is settled by how the score orders the candidate edits at each step, an order that evaluation by the final graph never observes. We measure it by replay on five discrete benchmark networks: at every step of searches that were actually run, every legal edit is rescored with the shipped implementation and against the ground-truth graph. On the three smaller networks the edit the search takes beats most accepted alternatives in oriented-edge F1, yet it is the best of them at only 45% of steps, and against its closest competitors in score it is truer 58% of the time. Reordering only the edits the score already accepts recovers 32.6% of the distance from a uniform draw to the single-edit ceiling, against 10.1% for the score's own order. Half of that gap is exact ties between Markov-equivalent successors, which the data cannot orient and the implementation breaks by enumeration order; a coin flip in its place moves the result significantly, in opposite directions on different networks. On the two larger networks the gap narrows to 4.3 points. Measured against what reordering can reach, a frontier language model in the reordering role recovers 76.2% on child and insurance, averaged over three runs, against 43.5% for the score's own order; its lead disappears when variable names are replaced by pseudonyms and survives plain-language descriptions. No open-weight model we tested beats the score significantly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.