Who Really Wins? Stress-Testing Symbolic Regression Benchmarks Across Metrics, Aggregations, and Datasets
Abstract
Symbolic regression (SR) benchmarks such as SRBench rank methods by averaging a single accuracy metric over repeated runs across evaluated datasets. Yet, the conclusions drawn from such leaderboards depend on implicit design choices, i.e., which error metric is reported, how repeated runs are aggregated, and which datasets are included. These design choices are rarely stress-tested. In this paper, we present a meta-study about the robustness of the SR ranking built on the recently proposed next-generation SRBench. Using each performance metric (MSE, MAE, , and training time) and each run-aggregation rule (mean, median, max, min, standard deviation, and four random single-run picks), we generate model rankings with a portfolio of multi-criteria decision-making (MCDM) schemes and introduce two robustness indicators: ranking-scheme robustness (sensitivity of rankings to both aggregation methodology and chosen metric) and dataset-composition robustness (sensitivity to correlation-aware dataset subsampling). Then, we perform metric-agnostic and aggregation-agnostic cross-consistency analysis. In addition, we complement our robustness analysis by isolating real-world (black-box) from physical-law-based (first-principles) datasets. Across SR methods and datasets, we find that no single method dominates every regime, but a small core—most consistently EQL and PS-Tree—remains robust across metrics, aggregations, and dataset compositions on real-world data, while QLattice and Operon rise to the top on first-principles problems. We validate five among the overall-top-ranked methods on five held-out external, well-established datasets. Our findings support the call for action of the SR community for a living benchmark and show that robustness indicators, rather than single averaged scores, should accompany future SR leaderboards. We publicly release all model scores and derived rankings as reusable benchmark artifacts, enabling new benchmarking studies without rerunning the original experiments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.