MAINFRAME-GYM: Benchmarking Agents on the Maintenance of Production Legacy Estates
Abstract
Coding agents are being deployed on COBOL estates just as the engineers who maintain them retire, yet no benchmark measures whether an agent can maintain a production mainframe system. Existing mainframe benchmarks test knowledge or single functions, and software benchmarks are mined from public repositories whose history and tests a production estate lacks. We introduce Mainframe-Gym, a benchmark of maintenance tickets on the 432k-line decommissioned billing estate of a large telecommunications company, and release the two pieces of infrastructure it required. Harwell runs the estate's COBOL and JCL batch chains, with DB2 emulated, off the mainframe. TaskForge is a human-AI collaboration harness that scales task construction: engineers direct the work, agents draft and verify the tasks, and the domain expert, who does not work in English, is consulted in their own language and reviews tasks in full, so scarce expert time goes to judgment rather than writing. Every grader is tested before any agent runs. We test five frontier models, each in its own or a common harness. The best pair solves five of eight tickets at least once, and two tickets defeat every pair. In a closer pilot study, agents usually find the right programs but make incomplete changes, missing the clone programs and archived versions that carry the same logic, and their own tests do not catch it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.