acceptodds
Under review as a conference paper at ICLR 2027

Small Steps on the Geodesic: Ability-Anchored Progressive Curricula for RL on Reasoning

Abstract

Curriculum design for RL with verifiable rewards (RLVR) is usually framed as a per-step problem: keep each batch maximally informative by admitting only prompts of intermediate pass rate, where the Bernoulli variance peaks. We argue that this objective is incomplete. Viewed as transport of the policy's ability distribution toward a target, training accumulates deviation along the whole path, which no per-step criterion can see. The direction of travel must be estimated from the model's own pass rates, and no estimate is exact: a step of length leaves the Wasserstein geodesic by an amount linear in , whereas covering the same displacement in re-anchored steps lets independent errors cancel, so the deviation grows only as . Bounded, repeatedly re-anchored increments therefore stay closer to the geodesic than greedy jumps to the currently best difficulty, which gives a formal reason why Krashen's "" hypothesis calls for small increments. Whether greedy filtering actually loses depends on how the estimation error scales with step length, an assumption we state but do not prove. The premise of the per-step view is nonetheless sound: under group-relative advantages, a prompt whose rollouts are all correct or all wrong contributes exactly zero gradient, and 61–62% of Hendrycks-MATH lies at one of these two extremes. We instantiate the analysis as Ri1A and compare it with ODF, AdaCuRL, DAPO, and DAPO with static difficulty ordering (DAPO+static). At equal wall-clock (five hours on Qwen3-4B with identical data, batch size and response limit), Ri1A outperforms ODF and AdaCuRL by 2.9 and 2.0 pp on 502 pooled competition problems (), at and lower cost per optimizer step. Against DAPO and DAPO+static trained on the full dataset for their complete schedules, Ri1A matches or exceeds their accuracy in 43–46% of the wall-clock time; its Hard-Avg margin over DAPO is +0.12/+0.79/+3.04 pp at 1.7B/4B/8B.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.