ALOER: Action-Level Orchestration of External Resources for Agentic Reinforcement Learning
Abstract
Agentic reinforcement learning (RL) trains large language models (LLMs) through environment interactions and requires resources beyond the training cluster, such as CPUs for code execution and GPUs for reward models. Existing service systems for agentic RL typically rely on static resource provisioning, which mismatches the multi-turn and intermittent pattern of external resource invocations, leading to severe inefficiency. We present ALOER, a unified system for elastic, action-level orchestration of external resources in agentic RL. ALOER utilizes a unified action-level formulation and an elastic scheduling algorithm to minimize action completion time (ACT) while satisfying heterogeneous resource constraints. Further, heterogeneous resource managers are tailored to efficiently support the action-level execution with elastic degree-of-parallelism on heterogeneous resources. Evaluation on well-established agentic RL tasks demonstrates that ALOER improves average ACT by up to 4.3, speeds up RL training steps by up to 1.5, and reduces external resource usage by up to 71.2.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.