Large Language Models as Generative Bayesian Policies
Abstract
Bayesian optimization is the dominant framework for sample-efficient scientific discovery, providing principled uncertainty estimates that guide expensive experiments. However, its reliance on hand-crafted, domain-specific representations limits broad applicability. Large language models (LLMs) offer flexibility as many experimental setups can be described in text, but they lack the calibrated uncertainty required in high-stakes scientific applications. Here we ask whether Bayesian optimization can supply the training signal for preference-based LLM alignment. We train a Gaussian process critic through deep kernel learning on the LLM's own embeddings. The same critic's acquisition function then provides a preference signal that trains the LLM to act as a generative Bayesian policy: candidates favored by the acquisition function become "preferred" completions for preference-based finetuning. The actor learns to generate high-utility candidates directly, thereby amortizing the acquisition search of standard Bayesian optimization into the language model itself. We evaluate on three scientific optimization benchmarks spanning reaction optimization, molecular and peptide design under a budget of 100 experiments. The actor's generation distribution shifts progressively toward higher-utility regions across iterations, and the trained model outperforms feature-based Bayesian optimization, LLM-guided optimization, and their union at different levels of integration. These results suggest that calibrated uncertainty can be the training signal that aligns language models with the decision-making demands of experimental discovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.