ViGa-SR: LLM-based Evolutionary Search for Symbolic Regression with Visual Priors Guidance from Multimodal Representation of Data
Abstract
Symbolic regression aims to derive concise formulas that both fit data and reveal the mechanisms underlying it, with evolutionary computation (EC) as the widely adopted approach. However, EC-based SR methods must navigate a vast search space to find suitable equations. Prior knowledge can ease this burden, but existing approaches either rely on single-modal information or apply guidance indiscriminately, neglecting rich multimodal insights that human experts naturally exploit to adjust their priors during evolution. We propose ViGa-SR, a framework that automatically discovers the underlying mathematical structure of data and exploits it as prior knowledge to guide LLM-based evolutionary search. The data is first interpolated and rendered through diverse visualization schemes, from which a multimodal LLM (MLLM) progressively identifies the potential substructures, operators, and coupling relationships among the variables in the data; each identified feature is then numerically verified, with its associated constants calibrated. The identification proceeds hierarchically. After each feature is verified, it is stripped from the data, and the MLLM re-examines the simplified data. This process iterates to expose progressively deeper substructures and terminates when no further regularity can be confirmed. The verified prior knowledge guides the subsequent LLM-based evolutionary search and is re-acquired with altered visualization schemes whenever evolution stagnates or population diversity drops. Compared with purely numerical feature extraction, MLLM-based feature extraction is more robust to noise and sample sparsity, and the use of visualizations as the communication medium keeps the entire process transparent and human-readable. On the Feynman dataset and LLM-SRBench, ViGa-SR matches or exceeds state-of-the-art performance, demonstrating strong feature-extraction capability and noise robustness.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.