AriadneBench: Evaluating Scientific Decisions by AI Agents in Cryo-EM Workflows
Abstract
Recent advances in AI agents are reshaping workflows for scientific discovery. However, reliable decision-making in cryo-electron microscopy (cryo-EM) remains challenging for agents, as they must align decisions with scientific objectives and autonomously adapt to multimodal evidence. We introduce AriadneBench, the first benchmark for evaluating these capabilities in real cryo-EM workflows. It comprises 180 tasks across eight task families and 13 workflows, targeting scientific goal alignment and adaptive autonomy. Workflow snapshots preserve the information available at each decision point, while expert annotations and auditable evidence tools enable joint evaluation of answer correctness, reference evidence coverage, and operational compliance. We evaluate seven state-of-the-art agent systems and find that the strongest achieves an overall Success Rate of only 63.21%. Agents perform better at configuring operations and recommending next steps than at particle selection and failure localization. Trajectory analysis shows that even complete reference evidence coverage does not guarantee a correct decision. To investigate whether distilled knowledge from our benchmark can improve these decisions, we induce one reusable skill per task family from training trajectories and expert annotations. On 37 held-out tasks, skill guidance improves Success Rate for five of seven agents, with an average gain of 4.25 percentage points across all seven agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.