When to Trust Your AI Co-Scientist? Learning Prior Reliability in Closed-Loop Bayesian Optimization
Abstract
AI “co-scientists” promise to accelerate discovery by proposing what to test next from knowledge distilled across millions of publications. Their reliability, however, depends on literature coverage: on under-studied targets they fabricate plausible candidates and fail silently, so a practitioner cannot tell a useful prior from a harmful one before committing costly experiments. Larger models mitigate but cannot eliminate failures where the literature is thin. We propose HARIBO, which learns a prior’s reliability from the experiments themselves. HARIBO adds the prior as a single feature of the Bayesian optimization (BO) surrogate. Its coefficient, fit to the sparse observations already being collected, provides an interpretable measure of trust that exploits, ignores, or inverts the prior as the data dictate. Across gene perturbation, protein variant effects, and antibody developability, and across language-model agents from 7B to frontier scale and protein foundation models, HARIBO improves on BO wherever the prior carries signal and matches it where the prior is noise. Under a systematically wrong prior, every scheme that fixes the prior’s weight in advance, by a constant, schedule, or model confidence, falls below BO, while HARIBO stays above it. The model is consulted before the first experiment and never again, which is enough: re-querying it every round instead, at 20–32× the requests and 124–234× the tokens, helps the resulting closed-loop agent on one of three domains and leaves it below knowledge-free BO on the other two. By learning trust from data rather than from model confidence, HARIBO lets AI co-scientists enter discovery loops safely.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.