EVALBENCH: Benchmarking Reference-Free PDE Scores from Ranking to Search
Abstract
PDE program search relies on reference-free scores, yet accurate ranking of a fixed candidate bank need not translate into accurate solutions when a score guides search. We introduce EVALBENCH, a benchmark that separates target identifiability, static score–error rank agreement (M1), and the error of the program returned by active search (M2). Core comparisons hold the proposer and fitter fixed; parameter, grammar, and proposer shifts test robustness. Across seven identifiable one-dimensional targets, five combined or resampled scores achieve median Spearman correlations of 0.813–0.817, compared with 0.417 for raw residual. Nevertheless, held-out scoring gives a median paired log10 error ratio of -0.0036 against baseline (95% interval [-0.032, 0.003]), showing no clear active-search gain. A Keller–Segel scale-family control demonstrates why ranking cannot resolve an incomplete public specification. SAAS, an operator-aware rank aggregator, fails its prespecified success criterion; a separate ten-operator development study localizes gains and persistent failures. We provide conditional transfer bounds under distribution coverage and simultaneous risk certificates for fixed selectors using independent calibration seeds from the deployment distribution. EVALBENCH makes specification ambiguity, ranking failure, and search limitations separately measurable, while reporting per-operator errors, catastrophic outcomes, coverage, and evaluation cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.