acceptodds
Under review as a conference paper at ICLR 2027

BoLT: A Benchmark to Democratize Black-box Optimization Research for Expensive LLM Tasks

Abstract

Optimization of LLM training and inference configurations, such as hyperparameters, data mixtures, prompts, and parallelism strategies, is critical to performance. However, it is often done heuristically, which can lead to suboptimal outcomes. These tasks are noisy, expensive, and derivative-free, making them a natural fit for sample-efficient Bayesian optimization (BO) and other black-box optimization (BBO) methods. This direction remains underexplored, largely because LLM training and inference are prohibitively expensive for most BBO researchers, and new methods are often evaluated only on synthetic test functions and small-scale datasets that fail to capture the challenges of LLM optimization. This impedes the development of BBO methods and makes their effectiveness on modern LLM tasks difficult to assess. We introduce BoLT, the first LLM-centric benchmark that democratizes LLM research for the BBO community. BoLT covers a broad set of well-motivated LLM optimization problems, involving multi-fidelity, multi-objective, heteroscedastic noise, high-dimensional search spaces, and black-box constraints. Each problem is grounded in thousands of real LLM experiments, and is fully reproducible and accessible through lightweight emulators or tabular datasets. We benchmark a wide range of BO and BBO methods on BoLT. BO methods generally outperform other baselines, but no single method wins across tasks, and methods remain far from optimal on some problems. These gaps underscore the need for benchmarks grounded in modern LLM tasks. An anonymized version of BoLT is available at https://github.com/anonom799/bolt.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.