RSIBench-Code: Can LLM Agents Handle Enterprise-Level LLM Development Code for Self-Improvement?
Abstract
Recursive self-improvement (RSI) aims to enable AI systems to contribute to the processes that produce stronger future models. A practical route is model self-training, but the code behind training, evaluation, serving, and infrastructure is still largely designed, maintained, and debugged by human engineers. This raises a central question: **can LLM agents handle the enterprise-level LLM-development code needed for self-improving AI**? We introduce **RSIBench-Code**, a benchmark for evaluating LLM agents on coding tasks reconstructed from practical LLM development problems in an enterprise. The benchmark contains 105 tasks across post-training, training infrastructure, serving, evaluation, and agent systems. These tasks cover implementation, repair, test writing, controlled experiments, and execution-path alignment. In each task, agents must locate and modify relevant code within repositories containing approximately 1.8 million source lines, with changing up to thousands of lines. RSIBench-Code supports long-horizon coding attempts with up to 150 agent turns and a 7,200-second solving budget. Experiments show that existing agents make meaningful but incomplete progress, suggesting that enterprise-level LLM-development code remains a key bottleneck for recursive self-improvement. Furthermore, RSIBench-Code provides a continuing source of executable evaluation tasks and potential training data for RSI in future work.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.