ProtoWorld: A Benchmark for World Modeling from Laboratory Protocols
Abstract
Laboratory protocols specify how to perform and reproduce experiments. Developing and executing them takes expertise and time, with some experimental workflows spanning months or even years. If AI agents could automate these procedures, they could save researchers time and effort. To do so, they would first need to model how experimental actions change the contents of laboratory containers, such as tubes. Thus, we introduce ProtoWorld, a benchmark for evaluating this aspect of world modeling from laboratory protocols. We give a model the instructions from the start of a protocol through a selected step and ask either how much liquid a container holds or which experimental substances (reagents) it contains after that step. If the volume cannot be determined from the instructions, the model must answer indeterminate. We report a ProtoWorld Full set of 100 manually annotated questions at 82 selected steps across 43 protocol records, including a ProtoWorld Hard subset of 40 questions manually labeled as hard. We use a large language model (LLM) to extract operations such as additions, transfers, and removals from the protocol. Then, we manually review these operations, and apply a deterministic program to these operations to calculate the reference answers. Across thirteen model configurations, the highest overall accuracy is 64.7% on ProtoWorld Full and 30.0% on ProtoWorld Hard. We hope ProtoWorld will help researchers evaluate and improve the world-modeling capabilities that agents need to design, refine, and execute laboratory protocols, ultimately helping accelerate scientific discovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.