acceptodds
Under review as a conference paper at ICLR 2027

Differentiate the Evaluator: Co-Searching Programs and Parameters with Exact Gradients

Abstract

Scientific models can be expressed as executable programs whose code represents mechanisms, equations, control flow, and numerical procedures. An AI agent may therefore propose the correct scientific mechanism yet have it rejected because its continuous constants are poorly calibrated. In program search, calibration is not merely post-processing: it helps determine which scientific ideas survive. Language models and gradient-based optimization offer complementary strengths, with language models proposing discrete program structure and gradients efficiently calibrating continuous constants. Existing approaches, however, either restrict candidates to fixed differentiable templates or require new tracing, compilation, or implementation work for each proposed program. We instead differentiate the evaluator once, allowing programs supplied as data to inherit exact reverse-mode gradients with respect to embedded constants. Our native differentiable virtual machine executes these programs at runtime and batches many parameter candidates through a single structural execution, making gradient calibration practical as the inner loop of AlphaEvolve-style evolutionary program search. We evaluate the method by co-searching nonlinear WENO shock-capturing schemes for computational fluid dynamics. From language-model-proposed programs, the system reliably recovers the expert-designed WENO-Z+ mechanism and recalibrates its parameters through full solver rollouts. The resulting schemes reduce error by 34–46% relative to WENO-Z on three unseen problems (32–41% against independent reference solutions) and, for one seed, by 26–47% relative to published WENO-Z+, while retaining fifth-order accuracy; a stronger shock test exposes a small undershoot that WENO-Z avoids. Ablations show that evolutionary search identifies the mechanism, while exact-gradient calibration produces most of the cross-problem improvement. Discrete structure also acts as a regularizer: given the same gradient freedom, a strong expert-designed family crosses a stability cliff and develops large shock oscillations, whereas every discovered program keeps a fixed exponent that stays clear of it. A second co-search over state-space programs exercises loops, branches, and matrix operations, and an accompanying optimizer study shows that gradients become increasingly important as parameter dimensionality grows. These results suggest that differentiating the evaluator rather than each candidate program enables AI-driven co-search over discrete scientific structure and continuous parameters without per-program differentiation or engineering effort.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.