Pareto Fronts for Symbolic Regression: Traversal Search towards the Frontier
Abstract
Symbolic regression (SR) seeks compact mathematical expressions that balance predictive accuracy and model complexity, but the combinatorial growth of the expression space makes effective search challenging. This work makes two contributions: (i) we find Pareto fronts for SR benchmarks extended up to length 21 and (ii) we leverage the results of the Pareto front generation to develop a structured traversal SR search framework over explicitly defined topology-conditioned symbolic landscapes. For the first contribution, we combine tree topology enumeration with capped sampling of node labels to construct longer expressions, extending empirical accuracy–complexity Pareto fronts to expression lengths up to 21. For the second contribution, we develop traversal search over expression states within tree topologies, using multistart n-mutation search and beam search to explore mutation neighborhoods. Our central novelty is an explicit formulation of expression states and mutation neighborhoods within tree topologies, enabling traversal strategies to be evaluated against reference ranks and empirical Pareto fronts while finding high quality expressions with a small fraction of the evaluations required for exhaustive search. On representative benchmarks, our approach identifies expressions close to the empirical reference fronts, achieving competitive predictive performance with lower complexity than many SR algorithms in SRBench. We perform systematic comparisons of traversal search configurations to provide insight into effective strategies for locating highly ranked expressions and the relationship between traversal search quality and evaluation efficiency. We further examine training–test alignment to complement the evaluation of traversal search performance. Together, these contributions extend empirical accuracy–complexity frontiers to larger expression spaces and show how the resulting topology-conditioned landscapes can be exploited for efficient traversal search. By establishing an empirical reference for what is attainable within the explored search space, we provide a concrete target for algorithm development, allowing algorithms to be evaluated by how effectively they approach known high quality solutions rather than only by comparison with existing algorithms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.