Beyond Fixed Groups: Adaptive Rollouts for Coding-Agent Reinforcement Learning
Abstract
Long-horizon coding-agent rollouts have highly variable latency: an agent may search a repository, edit files, and run several test suites before receiving one verifiable reward. Fixed-size GRPO allocates the same number of rollouts to every query, while round-based collection delays resampling and fresh work behind stragglers. We present ARCA, an event-driven realization of Adaptive Rollouts for Coding Agents. It processes each query when its current sampling stage completes, closes mixed groups, resamples all-wrong groups within a bounded budget, and refills fresh queries. The policy stays fixed during collection; the learner updates after collection and draining. In single-run experiments with Qwen3.6-27B on SWE-bench Verified, the schedule at fresh-query width records rollout steps/h versus for native GRPO (). Observed peak pass@1 on the 51-problem evaluation slice is versus , a gain of percentage points. At , reaches peak pass@1, while records the highest throughput, steps/h. These throughput comparisons use different mixed-group targets and do not measure total compute savings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.