acceptodds
Under review as a conference paper at ICLR 2027

AriadneBench: Evaluating Scientific Decisions by AI Agents in Cryo-EM Workflows

Abstract

Recent advances in AI agents are reshaping workflows for scientific discovery. However, reliable decision-making in cryo-electron microscopy (cryo-EM) remains challenging for agents, as they must align decisions with scientific objectives and autonomously adapt to multimodal evidence. We introduce AriadneBench, the first benchmark for evaluating these capabilities in real cryo-EM workflows. It comprises 180 tasks across eight task families and 13 workflows, targeting scientific goal alignment and adaptive autonomy. Workflow snapshots preserve the information available at each decision point, while expert annotations and auditable evidence tools enable joint evaluation of answer correctness, reference evidence coverage, and operational compliance. We evaluate seven state-of-the-art agent systems and find that the strongest achieves an overall Success Rate of only 63.21%. Agents perform better at configuring operations and recommending next steps than at particle selection and failure localization. Trajectory analysis shows that even complete reference evidence coverage does not guarantee a correct decision. To investigate whether distilled knowledge from our benchmark can improve these decisions, we induce one reusable skill per task family from training trajectories and expert annotations. On 37 held-out tasks, skill guidance improves Success Rate for five of seven agents, with an average gain of 4.25 percentage points across all seven agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.