acceptodds
Under review as a conference paper at ICLR 2027

Learning at and Beyond the Competence Boundary for On-Policy Post-Training

Abstract

On-policy post-training adapts learning to a model's current behavior, while adaptive curriculum learning adjusts task selection to its evolving competence. Yet adapting responses and selecting from existing tasks still leaves a critical gap: persistent failures reveal capability deficits but do not by themselves create the new practice needed to overcome them. To address this gap, we propose CASE, a Competence-Aware Selection with Error-targeted generation framework for adaptive on-policy post-training. CASE adapts both the response and task distributions: at each round, it retains tasks near the student's current competence boundary and uses a teacher agent to diagnose recurring failures and generate verified intermediate problems targeting the underlying weaknesses. The student then learns from freshly sampled responses on the selected and generated tasks using group relative policy optimization with executable rewards. We evaluate CASE with Qwen-3.5-9B across code, math, and finance, including analyses of task selection, error-targeted generation, and teacher choice. The results show complementary gains from selection and generation, with an average pass@1 improvement of 7.69 percentage points over the base model on out-of-domain benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.