CaliBO: Calibrating Incremental LLM Advice for Bayesian Optimization
Abstract
Large language models (LLMs) have shown promise for augmenting Bayesian optimization (BO), particularly when objective evaluations are scarce. When both an LLM and a Gaussian process (GP) provide objective predictions, a central question is how their means should influence acquisition. We introduce CaliBO, which calibrates LLM guidance by estimating its incremental predictive value beyond the current GP. Using predictions recorded before each evaluation, CaliBO learns whether the disagreement between the LLM and GP means predicts the GP’s subsequent residual error. A signed, shrunk coefficient produces a bounded mean correction, while GP posterior uncertainty remains unchanged for exploration. This allows CaliBO to follow, downweight, or reverse LLM advice according to its estimated incremental value, without requiring the LLM to outperform the GP as a standalone predictor. We bound acquisition perturbation and, under a zero-usefulness null and stated GP and noise-estimation assumptions, establish a sublinear advice-dependent term in the regret bound of an idealized GP-LCB variant. Across 49 hyperparameter-optimization tasks, CaliBO is competitive with ten baselines and significantly outperforms several of them. Ablation studies support the effectiveness of incremental calibration, while controlled experiments with synthetic advisors demonstrate benefits from signed correction under systematically inverted advice and reveal a failure case when uninformative advice still tracks the GP's error.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.