Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse
Abstract
Neuroscience data are highly fragmented across labs, experimental paradigms, and storage formats, and reuse often requires substantial manual effort to decipher each dataset's bespoke formatting choices and documentation. Agentic AI is a natural candidate to reduce this burden: LLMs read code and text quickly and with apparent attention to the low-level details humans tend to skim over. To measure how well agentic AI can perform data reuse, we select eight recent papers on large-scale mouse neural population recordings, covering diverse recording modalities and experimental paradigms. We provide coding agents with the data, code, and paper, and prompt them to reformat the data for a common downstream task: training a decoder from neural activity to task and behavioral variables. This task is a demanding benchmark target, since many conversion decisions are under-specified and several choices are defensible, yet a subtle but scientifically significant error can invalidate the resulting dataset in ways aggregate metrics do not reveal. We therefore evaluate agents using both outcome-based metrics and a process-based rubric scored independently by two human evaluators against manual reference solutions. We find that agents perform well on most subtasks, but rarely string together a fully error-free solution. We characterize the types of mistakes agents make, and suggest data-sharing best practices for agent-assisted data reuse. We further find that agents-as-judges are unreliable at catching errors on their own, but can automate process-based evaluation when given reference solutions. We show that newer models and prompt engineering improve performance, but no configuration we test eliminates end-to-end errors. All reference solutions, decision records, agent trajectories, and ratings are released as benchmark assets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.