acceptodds
Under review as a conference paper at ICLR 2027

EvoNeuroBench: Evaluating Coding Agents on Real Research Tasks in Neuroscience

Abstract

Whether large language model (LLM) agents can autonomously develop and iteratively improve neuroscience algorithms is an important question for their effective use in brain research. Successful code execution alone is insufficient because agents must also choose suitable methods, meet experimental requirements, and improve their solutions through feedback. However, these capabilities remain insufficiently evaluated across the diverse tasks encountered in neuroscience. To address this gap, we introduce EvoNeuroBench, a benchmark of 40 executable machine-learning tasks spanning multiple modalities, experimental settings, and biological scales. Each task is grounded in a research dataset, a published study or a public challenge and preserves the scientific objective and evaluation design of its source. Agents work for 12 hours in isolated environments with hidden deterministic evaluators, revising their solutions from quantitative feedback while every submission is retained, so that both final performance and progress through experimentation are observable. Using this framework, we compare five coding-agent configurations under matched conditions and contextualise their performance with published references. Even the strongest configuration reaches the published state of the art on only 25% of the tasks, and 58% of the tasks are reached by none. Success concentrates on problems that translate into standard supervised pipelines, whereas dense clinical segmentation, speech neuroprostheses and cross-site generalisation remain far from the reference. We find that some agents use thousands of submissions to infer evaluation labels from exact score feedback. Post-hoc auditing is therefore needed to distinguish feedback exploitation from genuine algorithmic progress. Together, EvoNeuroBench provides a unified, auditable testbed for how coding agents develop and improve neuroscience algorithms.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.