Dr.RL-Bench: Diagnosing Reinforcement Learning Runs with AI Agents
Abstract
Autonomous reinforcement-learning (RL) post-training requires agents to understand a training run before deciding how to improve it. Existing benchmarks primarily measure end-to-end model improvement, conflating training diagnosis with intervention design and experiment execution. We introduce Dr.RL-Bench, a benchmark that isolates diagnosis as an evidence-grounded reasoning task before intervention. Its 196 tasks are constructed from real GPU training runs and cover anomalies in data, training parameters, infrastructure, rewards, and environments. Agents freely investigate the provided code, configurations, logs, and generated responses, then report checkable Facts and evidence-supported Inferences. Expert-authored rubrics evaluate whether these findings identify relevant issues and explain their underlying causes. Experiments with eight LLM agents reveal substantial diagnostic gaps: the strongest model achieves 47.52/100, and all agents score lower on Inferences than on Facts. These results and case studies show how agents can recognize symptoms yet fail to connect dispersed evidence into a supported explanation. Dr.RL-Bench complements outcome-based evaluation by assessing what agents understand about a training run before they act.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.