acceptodds
Under review as a conference paper at ICLR 2027

Biological Context Does Not Reliably Improve LLM-Guided Experimental Optimization

Abstract

Autonomous labs increasingly deploy large language models (LLMs) as closed-loop biology optimizers that propose a design, read the outcome, then propose again. This is partly because of an assumption that biological knowledge helps an LLM choose better designs than Bayesian optimization, which has no access to biological context. We test that assumption on 45 biology tasks constructed from 2,485 experimental rows across 341 published papers, scored by oracles fitted to those rows, each run over 30 iterations under two prompt conditions. The domain-aware condition provides the published parameter names, units, and labels, and the domain-agnostic renames them so they are not biologically meaningful, with search space and feedback identical. Across 19 LLMs, 14 sit at or below a GP-UCB reference in both conditions, and biological context costs further ground, domain-agnostic outperforming domain-aware on 58.7% of tasks by trajectory and 61.9% by endpoint. The failure is specific to biology, reversing on education tasks built the same way. Domain-aware prompting tends to help models propose better initial designs, while making them respond worse to feedback. Part of the reason is anchoring, with models keeping the conventional categorical choice even when the oracle scores a different (published best) value higher. Post-training Llama 3.1 8B to frontier-level raw optimization does not transfer to biology-aware optimization improvements. Biology-aware optimization therefore looks like a separate capability that autonomous laboratory loops should target directly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.