acceptodds
Under review as a conference paper at ICLR 2027

Agentic Optimization for Asynchronous RL Post-Training

Abstract

Reinforcement learning (RL) post-training is computationally expensive, motivating asynchronous training for improved efficiency. Asynchronous RL, however, exposes a large, complex, and coupled configuration space across training, rollout, and resource orchestration. If configured improperly, asynchronous training will lose the efficiency advantages while introducing policy staleness. In this paper, we formulate a joint objective over model quality and training efficiency for asynchronous RL training and task LLM-based agents with iteratively optimizing it through hypothesis formulation, experimentation and evaluation. We evaluate both open- and closed-source agents across reasoning, multimodal and agentic post-training tasks using VeRL, a widely adopted open-source RL infrastructure. Our agents improve model quality and training throughput over human reference configurations while reducing per-task training GPU-hours by up to 50%, and distill transferable insights that accelerate optimization on a new task. Analysis of optimization trajectories provides insights into configuring asynchronous RL training and suggests directions for efficient RL post-training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.