acceptodds
Under review as a conference paper at ICLR 2027

PhysSymbolic-R1: Aligning Language Models for Physics-Aware Symbolic Regression

Abstract

Large Language Models (LLMs) have demonstrated strong capabilities in scientific reasoning, drawing growing attention to their potential in scientific discovery. However, scientific discovery fundamentally originates from deriving formal laws directly from observational data, known as Symbolic Regression (SR). This task poses a critical challenge to LLMs, stemming from an inherent tension between pre-trained LLMs’ proficiency in approximate reasoning, rooted in probabilistic text generation, and the high-precision demands of SR tasks. While recent methods employ complex external scaffolds to mitigate this limitation, such an iterative agentic paradigm remains computationally inefficient and fundamentally separates the model’s internal scientific knowledge from the symbolic regression process. To address this problem, we propose to directly equip LLMs with precise symbolic understanding and the ability to induce laws from raw observational data. In this paper, we introduce PhysSymbArena, a large-scale benchmark, comprising more than 160,000 diverse equations and 1.8B tokens of numerical symbolic data and physical description, to support post-training and extensive evaluation. Building on PhysSymbArena, we propose SymbolicLM, in which the symbolic regression capability of LLMs is systematically enhanced during post-training through mathematical and physical supervision. At inference time, we introduce a SymbolicSGA Refinement framework that uses quantitative fitting feedback to iteratively refine and recombine promising symbolic structures, enabling the model to correct structural errors and progressively recover more accurate governing equations. Experiments on multiple symbolic regression benchmarks demonstrate that SymbolicLM substantially outperforms representative symbolic regression methods in structural recovery while maintaining competitive numerical fitting performance. These results demonstrate that symbolic regression can be explicitly learned and strengthened as an intrinsic capability of LLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.