AutoCode-RL: Climbing the Difficulty Pyramid with Verifiable Synthetic Problems
Abstract
Reinforcement learning with verifiable rewards (RLVR) trains a coding agent by rewarding programs that pass tests, so the model learns only from problems it solves some of the time. Competitive-programming problems, however, form a pyramid: easy problems are plentiful and hard ones are rare, so the problems that can still teach a model become scarce as it improves.We propose AutoCode-RL, in which a frozen setter model adjusts the difficulty of existing contest problems for the solver being trained: plentiful lower-rated problems become harder, so that the solver sometimes fails, and problems beyond its reach become easier stepping stones toward them. An automated pipeline builds and validates tests for each rewrite, and the solver learns from whether its programs pass them.Training GPT-OSS-20B on rewrites of Codeforces problems released after the benchmark and base model raises its single-attempt solve rate (pass@1) on LiveCodeBench-Pro by a relative on easy and on medium problems. Our trained model also solves one problem from the hard split, where the base model fails all attempts. AutoCode-RL outperforms training on the original problems with the same budget and supervised fine-tuning on a stronger model's correct solutions. To our knowledge, the resulting scores are the highest reported for GPT-OSS-20B on LiveCodeBench-Pro.The simplified problems work as stepping stones: the trained model solves seven of their harder originals that the base model never solves, none of which appeared in training. The same method also improves Qwen3.5-9B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.