ResearchGym: Evaluating Language Model Agents on Real-World AI Research
Abstract
We introduce ResearchGym, a benchmark and execution environment for evaluating AI agents on end-to-end research. To instantiate this, we repurpose oral and spotlight papers from top ML/NLP conferences by preserving the datasets, evaluation harness, and baseline implementations but withholding the paper’s proposed method. This results in seven containerized execution environments comprising 54 gradable evaluation tasks. Within each environment, agents must propose novel hypotheses, run experiments, and attempt to surpass strong human baselines on the paper's metrics. Across bounded-time and extended evaluations of frontier LLM agents, agents improve over the provided baselines in only 57 of 235 evaluations and complete only 20% of tasks on average. We identify recurrent difficulties with time and resource management, confidence in weak hypotheses, and coordination of parallel experiments. Yet rare runs, agents match the performance of methods developed by expert researchers, indicating that frontier agents can occasionally reach state-of-the-art performance, but do so unreliably.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.