Pushing the Limits of Laboratory 3D Perception and Long-Horizon Planning via Protocol-Aligned Action Prediction
Abstract
AI Scientist systems are moving toward physical laboratory operation, which requires asset understanding, protocol alignment, and long-horizon state tracking beyond general scientific knowledge. We introduce LabHorizon, a training and evaluation framework for Protocol-Aligned Action Prediction. Level 1 predicts the next action from multi-view laboratory assets and protocol context; Level 2 generates parameterized action sequences from protocol windows and action pools. The release contains 6,400 samples: 3,000 training and 200 test samples per level, spanning 108 Level-1 assets and 2,342 unique wet lab protocols. Evaluations across major model families reveal persistent failures in asset-conditioned action selection, numerical parameters, action order, and variable dependencies. We train a Qwen3.6-35B-A3B domain model and build an Actor-Simulator-Selector agent. Training improves Level 1 accuracy from 47.5% to 63.5% and Level 2 final score from 25.34% to 41.00%; the agent further improves them to 66.5% and 45.32%. LabHorizon thus provides diagnostic evaluation and learnable supervision for laboratory action prediction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.