acceptodds
Under review as a conference paper at ICLR 2027

Do Learned Models of Dynamical Systems Answer Interventional Questions? A Benchmark with Oracle References

Abstract

Models of dynamical systems learned from data are used to predict the effects of interventions and to infer which variables drive which, yet they are evaluated mainly on forecasts or recovered equations. We pose five questions to every fitted model: a held-out forecast, a new initial condition, a state perturbation, a clamp do-intervention that replaces one variable’s equation, and the causal graph. Each trajectory and intervention score is placed between a trivial reference and an oracle that knows the true functional form and fits only its constants to the same data; graph scores are compared with the complete graph and with chance. Twelve methods that see only 105 noisy samples of one trajectory (sparse regression, Koopman operators, symbolic regression, neural ODEs and a conditional-independence test) are evaluated on all 63 ODEBench systems and on two new sparse-coupling suites, one generated under an analysis plan fixed before any method was run on it. Given the form, the data suffice at low noise: the trajectory-fit oracle generalises to a new initial condition on 98/88/63% of the predictable ODEBench systems at noise 0/1/5%, whereas the best learned method reaches 62% without noise and at most 37% with it; without noise, the vector fields of SINDy, Weak SINDy and PySR fit the training orbit to within 1% (median) yet miss the truth by 11–181% along the new trajectory. With noise, on the 82 systems with an absent edge, all with three to five variables, no learned method’s median error under a clamp at the end of the training window is detectably below 1, the error of predicting no response, while the oracle’s lies between 0.01 and 0.30. Graph F1 cannot separate any method from the complete graph on ODEBench, which scores 0.955; balanced accuracy and a threshold-free AUROC can, and AUROC exposes read-outs that rank edges well but threshold badly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.