Recursive Self-Improvement through Defect-Driven Verifiable Environments
Abstract
Recursive self-improvement (RSI) requires a model to identify what to learn next and obtain reliable supervision for capabilities it has not yet mastered. We introduce a defect-driven RSI framework that turns recurring model failures into verifiable learning environments. We find that many failures are prompt-recoverable: the model fails unaided but succeeds with defect-specific structural guidance. This observation motivates using unaided performance and assisted verification success to target capabilities that remain deficient but are accessible under guidance. We use initialized executable environments to adapt training-problem distributions, generate and verify guided solution trajectories, and train the learner without guidance in its inputs. After each round, we reassess the updated learner to adjust data allocation and use it for controlled rewriting of subsequent problem descriptions, closing the loop between capability improvement and curriculum adaptation. In experiments with SIRL-Qwen2.5-7B on three optimization modeling benchmarks, we obtain relative pass@1 gains of 34.02% on MAMO ComplexLP and 12.62% on OptMATH-bench-193 over the base model. Across rounds, we observe an overall increase in the proportion of unaided successes, with defect-specific guidance, we obtain times as many verified trajectories as with unaided generation under the same sampling budget. Together, we use defect-specific structural guidance to link learning-target selection, dynamic capability assessment, and training-data generation, and we use the constructed environments to support execution and verification throughout this recursive improvement process.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.