acceptodds
Under review as a conference paper at ICLR 2027

Dr.RL-Bench: Diagnosing Reinforcement Learning Runs with AI Agents

Abstract

Autonomous reinforcement-learning (RL) post-training requires agents to understand a training run before deciding how to improve it. Existing benchmarks primarily measure end-to-end model improvement, conflating training diagnosis with intervention design and experiment execution. We introduce Dr.RL-Bench, a benchmark that isolates diagnosis as an evidence-grounded reasoning task before intervention. Its 196 tasks are constructed from real GPU training runs and cover anomalies in data, training parameters, infrastructure, rewards, and environments. Agents freely investigate the provided code, configurations, logs, and generated responses, then report checkable Facts and evidence-supported Inferences. Expert-authored rubrics evaluate whether these findings identify relevant issues and explain their underlying causes. Experiments with eight LLM agents reveal substantial diagnostic gaps: the strongest model achieves 47.52/100, and all agents score lower on Inferences than on Facts. These results and case studies show how agents can recognize symptoms yet fail to connect dispersed evidence into a supported explanation. Dr.RL-Bench complements outcome-based evaluation by assessing what agents understand about a training run before they act.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.