acceptodds
Under review as a conference paper at ICLR 2027

Outside the Dictionary: LLM-Guided Equation Discovery Is Limited by What It Proposes

Abstract

Symbolic regression, including its recent language-model-guided variants, searches a fixed dictionary of primitives, and fails when the governing law needs a primitive outside that dictionary or a quantity that was never measured. Such failures are invisible to in-distribution evaluation: we prove that on a bounded observed domain, at every sample size and every positive noise level, some in-dictionary expression is statistically indistinguishable from the true law. We therefore score hypotheses by extrapolation, and use this to measure what limits a standard discovery loop in which a language model proposes closed-form laws that are then fitted, selected, and tested by active probing. We contribute Crisis-60, 60 procedurally generated environments in which a low-order polynomial fits the observed data almost perfectly but a law counts as recovered only if it extrapolates to a disjoint region with R^2 > 0.999; its control category, whose law is itself a polynomial, is solved at 100.0% by polynomial regression against 0.0% elsewhere. We also contribute the Epicycle Index, a ratio of extrapolation error to description length anchored at the incumbent theory, and a decomposition of recovery into proposal coverage, whether a correct law was ever proposed, and selection efficiency, whether it was then chosen. Selection is nearly saturated: it returns a correct law on 98.0% of the instances where one was proposed, so a perfect selector would gain only 1.1 points, whereas on 3 of 15 structural families no correct law is ever proposed. This accounts for why the loop's parsimony penalty and active probing contribute nothing measurable (95% upper limits 0.6 and 6.4 points), and the same limit holds for LLM-SR (Shojaee et al., 2025), whose representation and search differ from ours. Under observational noise the limit moves to selection, whose efficiency falls to 14.3%. The loop recovers 53.3% of laws (95% CI 42.8 to 63.9) against 23.9% for genetic symbolic regression and 31.1% for LLM-SR at matched budget; with five times the candidates LLM-SR reaches 51.7%, so the loop's advantage is efficiency rather than attainable recovery. Because the decomposition needs only the candidates a search already scores, it tells the builder of any such system which half is worth improving: here, the proposal distribution. Code, benchmark, and all logged runs are available at https://anonymous.4open.science/r/outside-the-basis-1867.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.