Soft-Rejection Informed Bayesian Optimization for Expensive Sequential Experiments
Abstract
Bayesian optimization (BO) is widely used for expensive black-box optimization, but in many settings each evaluation is itself a long sequential experiment whose final outcome is the optimization target. Early stopping can further reduce cost; however, existing methods inherit the reliability of an underlying surrogate predictor, and can fail when trajectories cross, i.e., when a configuration that looks poor early overtakes others by the endpoint. We formally define this early stopping problem for sequential black-box optimization, and propose Soft-Rejection Informed Bayesian Optimization (SRBO) to free the stopping rule from this dependence. SRBO is built around an endpoint forecaster, a Direct-Forecasting Gaussian Process (DF-GP), one GP per decision step, that maps the partial trajectory and configuration directly to the endpoint. This forecaster is used in two ways. A Calibrated Composite Predictive Sequential Test (CCPST) turns the DF-GP forecast into a stopping rule calibrated by held-out cross-validation: for the cross-validated rule, under a per-epoch independence condition on the held-out stop indicators, the rate of falsely stopped winners has a finite-sample bound that holds even when trajectories cross and the forecast is mis-specified, and for the deployed rule it is audited on every benchmark. The same forecast is also fed back into the BO surrogate as a pseudo-observation when a trial is stopped to learn from rejected configurations. On eight cost-matched surrogate benchmarks built from real experimental and learning-curve data, SRBO attains the best average rank of the seven methods run on all eight and, against up to ten baselines, the lowest mean final regret on three of the five experimental benchmarks, significantly against every baseline on each: on in vitro permeation testing (IVPT) it is below the strongest baseline and to below the others (), on UCI Concrete compressive strength below the strongest (), and on a methane-coupling catalyst screen below it (). On the two remaining experimental benchmarks SRBO shows no significant difference from the leaders on one and trails single-fidelity BO on the other (), and on three learning-curve tasks BOHB leads.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.