acceptodds
Under review as a conference paper at ICLR 2027

Outcome-Ranked Proposal Tuning: Training LLMs to Transfer Experience Across Bayesian Optimization Tasks

Abstract

Bayesian optimization (BO) is widely used for expensive black-box optimization. When related optimization problems arise repeatedly, experience from previous tasks can improve optimization on new ones. We aim to train a large language model (LLM) to transfer the knowledge from the prior tasks to new ones through BO initialization. We introduce Outcome-Ranked Proposal Tuning (ORPT), which trains an LLM to propose initialization points based on how they affect the subsequent BO process. ORPT trains the LLM to favor initialization candidates based on their effect on subsequent BO, rather than their standalone objective values. We further mathematically show that these learned preferences are linearly connected to the expected outcome of downstream BO. Experiments on synthetic Branin optimization and antimicrobial peptide design show that ORPT outperforms a broad range of single-task, multi-task, and LLM-based BO baselines, yielding stronger initializations and better downstream optimization. These gains persist through subsequent BO, while ablations confirm that downstream BO outcomes provide a stronger training signal than candidate objective values alone. Overall, our results show that learned initialization is an effective mechanism for transferring experience across related optimization tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.