Learning from Successes and Failures: Efficient Self-Evolution through Asymmetric Supervision
Abstract
Reasoning models can continue to improve after instruction tuning, yet existing approaches often either use reinforcement learning for self-evolution or rely on pre-existing per-question data for supervised refinement. We ask whether an instruction-tuned model can improve by directly learning from the structure of its own generated experience. Our key insight is asymmetric supervision: verified successes provide constructive targets for SFT, while competing incorrect answers provide comparative signals without treating failed reasoning trajectories as negative demonstrations. Starting only from a high-level domain prior and difficulty specifications, our framework constructs capability-matched curricula from model-generated problems and converts verified experience into success amplification and answer-level failure discrimination, using parameter-efficient LoRA updates on a compact training set. On Qwen3-4B and Qwen3-8B, verified-success SFT alone improves average accuracy across five challenging mathematical benchmarks by 2.27 and 2.28 percentage points, demonstrating that self-improvement without reinforcement learning is possible. Failure-aware supervision provides further gains, yielding total improvements of 3.12 and 2.97 percentage points, respectively, with aggregate gains across a broader nine-benchmark mathematical reasoning suite. Repeated-sampling analyses further indicate improved access to successful solutions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.