acceptodds
Under review as a conference paper at ICLR 2027

COBRA: Completion-Order Batching as a Reward-stratified Advantage

Abstract

COBRA improves early learning from verifiable training problems in single-response reinforcement learning. We retain every rollout,form update blocks by completion order, and switch from local to pooled reward moments after an initial learning phase. From a sharedformat-calibrated Qwen2.5-7B start, two runs improve update-80 accuracy over random batching by 3.15–3.49 percentage points onDeepMath and 4.20–7.18 on MATH500. Each arm has seen about 45,900 distinct training problems in 61,440 prompt presentations. On theretained evaluation grid, mean MATH500 accuracy first exceeds 70% at update 80 versus 120–160 for random/local, after 24–35% fewerdistinct training problems. Completion grouping creates length and reward strata; an exploratory, single-seed centering interventionsupports local centering as a contributor to subsequent response growth. Switching to pooled moments limits that growth through 200updates on Qwen2.5-7B and Qwen3-4B. On held-out three-to-seven-person logic puzzles, the same recipe gains 2.35–6.13 percentagepoints over scheduled random grouping.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.