acceptodds
Under review as a conference paper at ICLR 2027

EVALBENCH: Benchmarking Reference-Free PDE Scores from Ranking to Search

Abstract

PDE program search relies on reference-free scores, yet accurate ranking of a fixed candidate bank need not translate into accurate solutions when a score guides search. We introduce EVALBENCH, a benchmark that separates target identifiability, static score–error rank agreement (M1), and the error of the program returned by active search (M2). Core comparisons hold the proposer and fitter fixed; parameter, grammar, and proposer shifts test robustness. Across seven identifiable one-dimensional targets, five combined or resampled scores achieve median Spearman correlations of 0.813–0.817, compared with 0.417 for raw residual. Nevertheless, held-out scoring gives a median paired log10 error ratio of -0.0036 against baseline (95% interval [-0.032, 0.003]), showing no clear active-search gain. A Keller–Segel scale-family control demonstrates why ranking cannot resolve an incomplete public specification. SAAS, an operator-aware rank aggregator, fails its prespecified success criterion; a separate ten-operator development study localizes gains and persistent failures. We provide conditional transfer bounds under distribution coverage and simultaneous risk certificates for fixed selectors using independent calibration seeds from the deployment distribution. EVALBENCH makes specification ambiguity, ranking failure, and search limitations separately measurable, while reporting per-operator errors, catastrophic outcomes, coverage, and evaluation cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.