acceptodds
Under review as a conference paper at ICLR 2027

MAINFRAME-GYM: Benchmarking Agents on the Maintenance of Production Legacy Estates

Abstract

Coding agents are being deployed on COBOL estates just as the engineers who maintain them retire, yet no benchmark measures whether an agent can maintain a production mainframe system. Existing mainframe benchmarks test knowledge or single functions, and software benchmarks are mined from public repositories whose history and tests a production estate lacks. We introduce Mainframe-Gym, a benchmark of maintenance tickets on the 432k-line decommissioned billing estate of a large telecommunications company, and release the two pieces of infrastructure it required. Harwell runs the estate's COBOL and JCL batch chains, with DB2 emulated, off the mainframe. TaskForge is a human-AI collaboration harness that scales task construction: engineers direct the work, agents draft and verify the tasks, and the domain expert, who does not work in English, is consulted in their own language and reviews tasks in full, so scarce expert time goes to judgment rather than writing. Every grader is tested before any agent runs. We test five frontier models, each in its own or a common harness. The best pair solves five of eight tickets at least once, and two tickets defeat every pair. In a closer pilot study, agents usually find the right programs but make incomplete changes, missing the clone programs and archived versions that carry the same logic, and their own tests do not catch it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.