acceptodds
Under review as a conference paper at ICLR 2027

Learning from Successes and Failures: Efficient Self-Evolution through Asymmetric Supervision

Abstract

Reasoning models can continue to improve after instruction tuning, yet existing approaches often either use reinforcement learning for self-evolution or rely on pre-existing per-question data for supervised refinement. We ask whether an instruction-tuned model can improve by directly learning from the structure of its own generated experience. Our key insight is asymmetric supervision: verified successes provide constructive targets for SFT, while competing incorrect answers provide comparative signals without treating failed reasoning trajectories as negative demonstrations. Starting only from a high-level domain prior and difficulty specifications, our framework constructs capability-matched curricula from model-generated problems and converts verified experience into success amplification and answer-level failure discrimination, using parameter-efficient LoRA updates on a compact training set. On Qwen3-4B and Qwen3-8B, verified-success SFT alone improves average accuracy across five challenging mathematical benchmarks by 2.27 and 2.28 percentage points, demonstrating that self-improvement without reinforcement learning is possible. Failure-aware supervision provides further gains, yielding total improvements of 3.12 and 2.97 percentage points, respectively, with aggregate gains across a broader nine-benchmark mathematical reasoning suite. Repeated-sampling analyses further indicate improved access to successful solutions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.