acceptodds
Under review as a conference paper at ICLR 2027

WHEN QUEUE POLICY BECOMES LEARNING POLICY: AUDITING BATCH PATHS IN AGENTIC RL

Abstract

An asynchronous rollout queue can change learning even when every trajectory is eventually consumed: it determines which complete prompt groups share an optimizer step and in what order. We audit this batch-path effect using immutable recorded experience and persistent-Adam GRPO. In a prospective six-seed Qwen3-0.6B experiment, reversing two task-family blocks changes final product-minus-shared success by -23.6 percentage points [-30.5, -16.3], with experience, tokens, initialization, and update count matched. Each order favors its first family relative to balanced batches. An independent six-seed intervention clears only Adam’s first moment at the family switch, attenuating the contrast by +21.88 points [+16.54, +27.21]; all 168 checkpoint-and-optimizer control pairs are exact. The reset mainly removes the first-family advantage and reduces aggregate accuracy. We make the batch-composition choice explicit through a per-update contract for required coverage and histogram total variation, with earliest-feasible complete-group admission. Simple per-family queues exactly match the general allocator in the tested two-family case. Measured multi-turn traces expose composition violations and an explicit waiting cost. A separate 3B learning result remains unconfirmed, and a public-code pilot fails validity and its numerical screens. The evidence supports a controlled queue-path diagnosis and a composition guarantee, without establishing broad learning generality or time-to-quality improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.