SemanticBBOBench: A Benchmark for Semantic Black-Box Optimization with Large Language Models
Abstract
Black-box optimization is central to many scientific and engineering disciplines. While numerical methods such as Bayesian optimization (BO) can successfully optimize black-box problems, these methods struggle to incorporate the semantic context that often accompanies real-world experimental problems, including domain knowledge, prior experiments, and expert intuition. Using this information could substantially reduce experimental costs and improve outcomes. Large language models (LLMs) can reason over both semantic and numerical data, offering a pathway for bringing this context into the optimization loop. We present SemanticBBOBench, a black-box optimization benchmark designed for testing LLM-based methods with semantic context. It provides a uniform framework containing 42 optimization problems spanning diverse scientific and engineering domains, each paired with a curated context at multiple controlled levels of detail and types of information. Our experiments show that semantic context provides its greatest benefit early in the optimization process, enabling substantially stronger initial configurations when few numerical observations are available. We also show that misleading semantic information can persistently bias LLM-based search. These results highlight both the potential and limitations of LLM-based black-box optimization and establish SemanticBBOBench as a framework for developing methods that use such information effectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.