acceptodds
Under review as a conference paper at ICLR 2027

ResearchGym: Evaluating Language Model Agents on Real-World AI Research

Abstract

We introduce ResearchGym, a benchmark and execution environment for evaluating AI agents on end-to-end research. To instantiate this, we repurpose oral and spotlight papers from top ML/NLP conferences by preserving the datasets, evaluation harness, and baseline implementations but withholding the paper’s proposed method. This results in seven containerized execution environments comprising 54 gradable evaluation tasks. Within each environment, agents must propose novel hypotheses, run experiments, and attempt to surpass strong human baselines on the paper's metrics. Across bounded-time and extended evaluations of frontier LLM agents, agents improve over the provided baselines in only 57 of 235 evaluations and complete only 20% of tasks on average. We identify recurrent difficulties with time and resource management, confidence in weak hypotheses, and coordination of parallel experiments. Yet rare runs, agents match the performance of methods developed by expert researchers, indicating that frontier agents can occasionally reach state-of-the-art performance, but do so unreliably.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.