Beyond Motion Execution: Benchmarking Embodied Agents on Wet-Lab Operations and Procedures
Abstract
Completing a manipulation sequence does not guarantee wet-lab task success: an agent may still spill a reagent, miss a target mass, or overshoot a titration endpoint. We therefore introduce a 20-task dual-arm simulation benchmark, called WET-OP Bench, which couples manipulation with evolving experimental states. Tasks are organized around laboratory operations composed of shared atomic manipulation skills, spanning basic operations and experimental procedures of varying horizons. Task-specific demonstrations are provided for fourteen tasks, while six are held out to assess zero-shot generalization to new operation compositions and experimental conditions. For demonstrated tasks, controlled variations in task configuration, background, and instruction enable systematic evaluation of within-task generalization. The simulator couples manipulation with dynamic liquid transfer, quantitative instrument feedback, and procedure-conditioned appearance changes; the same underlying states drive observable feedback and success checks for containment, quantitative targets, and reaction outcomes. Reusable interaction graphs implement these state–action couplings across tasks, while Lab2Sim, an agentic asset-compilation pipeline, constructs specification-grounded, validated manipulation-ready assets. Finally, we conduct a capability-oriented evaluation of representative vision-language-action models and harness-based agents under controlled within-task and zero-shot settings, combining task success rates with skill-aligned stage scores for partial-progress tracking and failure diagnosis. The project website is available at https://beyond-motion-execution.pages.dev.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.