Evidence-Grounded Bayesian Optimization with Collective LLM Reasoning
Abstract
Expensive black-box optimization is challenging under tight evaluation budgets, particularly when reliable prior knowledge is limited. In such settings, context-informed warm starts can be ineffective or misleading. We propose a Bayesian optimization (BO) framework that grounds collective large language model (LLM) reasoning in experimental evidence. The framework uses design of experiments (DoE) to guide evaluations of the target system and derive statistical evidence that complements available prior information. Multiple LLM agents reason over this shared evidence, and their aggregated proposals guide search-region refinement and sequential BO. Our theoretical analysis establishes robustness properties of proposal aggregation, while experiments show that aggregated proposals achieve a lower mean normalized distance to reference optima than individual-agent proposals. On complex yet structured multimodal benchmarks with limited prior knowledge, our method reaches stringent regret targets with 8.8%–70% fewer evaluations than Gaussian process (GP)-based and LLM-assisted baselines. In FCNet hyperparameter-tuning ablations, evidence-based variants often reach top 0.5% configurations within 38 evaluations, earlier than non-evidence-based variants. These results support combining structured experimental evidence with collective LLM reasoning to improve the evaluation efficiency of black-box optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.