acceptodds
Under review as a conference paper at ICLR 2027

A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization

Abstract

Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making. However, existing Agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We introduce AgenticBBO-Bench, a cross-domain benchmark for Agentic BBO spanning synthetic optimization, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol. Agentic BBO achieves higher family-averaged scores than direct candidate generation by the same LLM in all five domains and outperforms the best evaluated numerical optimizers in four. We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search. Our results show that additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable, and numerical optimizers can sometimes continue effectively from search trajectories established by the agent. Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the same Codex agent harness. GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models. AgenticBBO-Bench provides a common standardized leaderboard for comparing future general-purpose models and agent systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.