LLM-PySR: Large Language Model-Guided Symbolic Regression via Search Control
Abstract
Scientific equation discovery must combine broad domain priors with strict numerical testing. Symbolic regression supplies numerical grounding but faces a combinatorial search space, whereas many language-model systems ask the model to propose or select formulas directly.We systematically test how alternative divisions of decision-making authority between language models and numerical symbolic regression affect equation discovery. We compare role specifications in which the language model acts as equation author, candidate decider or search controller, alongside end-to-end language-model and purely numerical baselines. In the controller setting we propose here, implemented as LLM-PySR, language models specify variables, operators, transformations and search depth; symbolic regression enumerates and fits expressions; and fixed rules with numerical safeguards govern retention. Across 74 AI-Feynman equations and seven complex formula-recovery tasks, search control achieved the strongest observed balance of accuracy, complexity, stability and cost. Notably, on an independent battery dataset, LLM-PySR identified a compact piecewise-linear relation between early voltage-curve displacement and cycle life. Overall, LLM-PySR provides a token-efficient, interpretable, and empirically effective approach for scientific equation discovery, bridging the gap between LLM guidance and classical symbolic regression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.