One Static Schedule Does Not Fit All: Platform-Aware and Runtime-Adaptive Scheduling for Agentic RL Rollouts
Abstract
Agentic reinforcement learning (RL) is increasingly important for advancing large language models, yet synchronous rollouts suffer from severe long-tail latency and poor resource utilization. Existing systems mitigate this bottleneck with predict-then-schedule policies that estimate trajectory durations and prioritize likely stragglers. However, two challenges limit their effectiveness. First, the optimal schedule is platform-dependent: LLM request latency depends on continuous-batch composition, while tool operations contend for parallel executors. Second, token-count and tool-latency predictions inevitably deviate from realized execution, potentially making a static schedule ineffective or even harmful. We present STARS, a simulator-driven framework for platform-aware and adaptive rollout scheduling. Positioned between the agentic RL framework and execution engines, STARS controls LLM admission and tool dispatch without modifying the training algorithm or underlying engines. STARS profiles the target LLM engine and tool pools offline, then uses simulation-based planning to evaluate the counterfactual effects of candidate scheduling decisions and generate a platform-specialized ranked plan for each rollout batch. The plan accounts for composition-dependent LLM latency and tool-executor parallelism. At runtime, an impact-triggered repair mechanism detects consequential deviations and incrementally reschedules the uncommitted portion. STARS exposes a policy-agnostic interface while providing a concrete simulator-guided policy. Experiments across diverse platform configurations demonstrate an average 1.12X speedup in rollout throughput over state-of-the-art synchronous RL systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.