DIGEST: Discovery-Guided LLM Tuning for Symbolic Regression
Abstract
Symbolic regression (SR) seeks to discover mathematical expressions directly from data, without requiring task-specific scientific descriptions. Despite tremendous progress in both conventional search-based methods and recent LLM-based methods, two critical limitations largely remain. First (Limited Generation Flexibility), conventional search-based methods rely on either fixed search rules or specialized learned generators, severely limiting adaptive candidate generation across diverse expression structures. Second (Underutilized Search Discoveries), most LLM-based methods cannot reliably recover ground-truth expressions because they fall short in effectively leveraging the LLM itself to thoroughly digest past expressions or adaptively guide future search. To address these limitations, we propose DIGEST. To render greater flexibility in generation, DIGEST repurposes a pretrained decoder block from a lightweight LLM into a symbolic policy for Genetic Programming (GP) population seeding, thereby leveraging the broader modeling capacity of the LLM. To more thoroughly utilize search discoveries, high-quality expressions discovered by GP are fed back to update the LLM-derived policy online, thereby enabling the LLM to progressively internalize past search discoveries and translate them into improved subsequent population generation. Experiments show that on two ground-truth benchmarks, DIGEST outperforms the strongest competitor by up to 2.59% in recovery. Among LLM-based baselines, it improves recovery by at least 54.91% while using at most 17.4% of the strongest competitor's mean search time. On 122 BlackBox datasets, DIGEST achieves the highest accuracy while maintaining a moderate model complexity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.