acceptodds
Under review as a conference paper at ICLR 2027

Do Candidate-Pool Benchmarks Faithfully Evaluate LLM Optimizers?

Abstract

A benchmark for a generative optimizer should evaluate the designs that the optimizer proposes. In scientific optimization, however, new measurements are often unavailable, and an optimizer is instead scored by Selecting from a pool of measured candidates or by Mapping its proposals to the nearest stored design. We ask whether these candidate-pool protocols preserve conclusions about prospective optimization. Across four materials benchmarks, larger search steps by large language model (LLM) optimizers are associated with fewer evaluations to reach the target region. In a controlled prompt experiment, a one-line instruction changes this reaching time without changing the underlying model. These observations motivate the trajectory-imprinting hypothesis: a pool can inherit geometric structure from the optimization runs that produced it. To test fidelity beyond the pool, we directly evaluate optimizer-proposed designs using physics-based simulations of planar MOSFETs and GaAs deposition. In GaAs deposition, Gaussian-process (GP) optimizers reach lower nonuniformity under Generation, whereas LLM optimizers reach the pool optimum first under Selection. Direct evaluation of mapped proposals reveals changes in measured-and-feasible status in 38.4% of MOSFET and 19.0% of deposition pairs. Among pairs feasible on both sides, mean projection credit is larger for GP optimizers than for LLM optimizers. We construct core-set pools from designs collected during previous optimization runs, keeping the initial design and the best candidate fixed. These pools have different measured geometric properties, but optimizer orderings under Selection still differ from those under Generation. Proxy fidelity must therefore be assessed for the optimizer set, pool, protocol, and metric rather than assumed from pool geometry alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.