Can LLMs Design Experiments? Benchmarking Experiment Design for Autonomous Labs
Abstract
Large language models (LLMs) have demonstrated strong capabilities in chemical reasoning and reaction prediction. However, knowing what reaction to perform does not necessarily imply knowing how to perform it in the laboratory. Experimental design requires translating scientific knowledge into concrete decisions about materials, operations, conditions, equipment, and execution order. Yet this capability remains largely underexplored in current LLMs. We introduce LabProtocol, a dataset and benchmark for learning and evaluating experimental design. LabProtocol aligns chemical reactions with real experimental procedures and represents each procedure as a sequence of structured, parameterized laboratory operations. These operations capture the materials, conditions, equipment, and temporal dependencies required to perform an experiment. Using LabProtocol, we evaluate leading LLMs on tasks ranging from complete procedure generation to single-operation completion. We find that current LLMs struggle to generate complete experimental procedures. More surprisingly, they still perform poorly when asked to predict only a single missing operation, suggesting that the limitation goes beyond long-horizon planning and reflects a lack of experimental procedural knowledge. We further fine-tune Qwen on LabProtocol and observe modest improvements in complete protocol generation but little improvement in single-operation completion. This contrast suggests that learning global protocol patterns does not necessarily yield fine-grained procedural understanding. LabProtocol provides a foundation for developing autonomous laboratory agents that translate scientific objectives into executable experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.