OpenSIR: Open-Ended Self-Improving Reasoner
Abstract
Self-play promises to advance LLM reasoning without human-annotated data, yet existing methods lack open-ended learning and yield only marginal or even negative gains on post-trained models. We introduce Open-Ended Self-Improving Reasoner (OpenSIR), a self-play framework where a single LLM alternates teacher and student roles to generate and solve novel problems without external verifiers or annotated data. From a single seed, OpenSIR sustains open-ended exploration via diversity rewards pushing toward unfamiliar concepts and difficulty calibration keeping problems learnable. Across seven math benchmarks, OpenSIR consistently improves all models, averaging +3.6 on instruction models and +3.1 on reasoning models, outperforming recent self-play baselines. Without any annotated data, it surpasses GRPO baselines trained on over 7K examples. Despite training only on self-generated math, OpenSIR is the only self-play method that transfers to general reasoning, improving reasoning models by at least +4.4 points across three benchmarks. Our analysis shows that, with diversity rewards, generated problems progress from basic arithmetic and algebra to advanced topics such as calculus and optimisation, reaching regions of the problem space beyond human-curated datasets. Without diversity rewards, problems concentrate in a narrow region of this space and accuracy falls on both mathematical and general reasoning. Difficulty calibration depends on teacher-student co-evolution: freezing the teacher destabilises problem difficulty and lowers accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.