When Not to Trust a Global Gaussian Process Posterior: Local Scoring and Progressive Dissonance for Multi-Modal Scientific Discovery
Abstract
Gaussian-process Bayesian optimization (GP-BO) is a standard approach to budget-constrained black-box optimization, but its acquisition decisions rely on the quality of a global GP posterior. We study when this reliance becomes unreliable in scientific discovery problems with phase boundaries, feasibility gaps, threshold-derived responses, and disconnected high-value regions. In these settings, a globally smooth posterior can be miscalibrated precisely in high-response regions: on LGPS, nominal % GP coverage drops from overall to in the top response decile. We introduce two alternatives that reduce dependence on global posterior ranking: ProSe, which combines local -nearest-neighbor scoring with spatial anti-clustering, and ATLAS, which searches for progressive disagreement with a hierarchy of simple models while probing spatial frontiers. Our results show a regime split rather than universal superiority. On a smooth synthetic control, several GP methods and random search attain Recall@100 — the fraction of distinct target peaks recovered within 100 evaluations — of . In contrast, on a strict LGPS compositional-family holdout, ATLAS achieves Recall@50/100 of and reaches successful discoveries in a median of evaluations versus for MAP-Elites, though its advantage over MAP-Elites is directional and borderline after Holm correction (). On discriminative steel search, ATLAS and SAGE — a -NN-scored baseline that anchors part of its candidates on observed points — each significantly outperform eight of ten alternative GP-based acquisition, quality-diversity, and randomized-search baselines after Holm correction. Finally, an audit of five look-up benchmarks constructed from public materials datasets finds that four show no separation among methods under our protocol, due to saturation, positional-encoding collapse, or random-search matching. These results suggest using global posterior acquisition on smooth, calibrated landscapes, and favoring locality, dissonance, and diversity when its modeling assumptions are violated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.