acceptodds
Under review as a conference paper at ICLR 2027

Multi-Dataset Symbolic Regression: Benchmark Construction and Reference Methods

Abstract

Most existing symbolic regression methods focus on discovering equations from a single dataset, whereas scientific experiments often require identifying a shared symbolic form from multiple datasets collected under different experimental configurations. To address this gap, we formulate the multi-dataset symbolic regression (MDSR) problem as the discovery of shared functional forms across datasets generated under varying experimental conditions, such as different parameter settings. We construct a benchmark suite for MDSR, comprising a real-world experimental benchmark from three underdamped oscillators and a synthetic benchmark generated from 59 transformed physics-derived problems under varying conditions. Unlike conventional symbolic regression benchmarks, our synthetic benchmark includes test datasets with unseen parameter values, controlled Wasserstein shifts in parameter distributions, and varying noise levels, enabling systematic evaluation of how shared functional forms discovered from training sets generalize under parameter-distribution shifts and observation noise. We further provide two reference solvers, a PySR-based Multi-stage Selection Symbolic Regression method (MSSR) and a Large-Language-Model-based Multi-Dataset Symbolic Regression method (LLM-MDSR), adapted from LLM-SR to support multiple datasets. Experimental results show that MSSR can rediscover shared functional forms equivalent to the theoretical formula on the real-world benchmark. On the synthetic benchmark, without specified parameters, MSSR identified shared functional forms with \(R^2>0.9999\) for 48 of 59 benchmark problems (81.3 %), while LLM-MDSR discovered shared functional forms with \(R^2>0.9999\) for 52 of 59 benchmark problems (88.1 %), on the training set.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.