TreeRollout: Resource-Aware Tree Sampling for Efficient Agentic LLM Post-Training
Abstract
Generating rollouts is a major cost in group-based reinforcement learning and rejection sampling for large language models. In agentic tasks, each trajectory alternates between GPU decoding and tool execution. The independence of individual trajectories leads to continuously fluctuating demand for concurrent tool execution and GPU decoding, intermittently leaving GPUs or tool executors idle. Scheduling can improve resource utilization, but spare capacity can remain unused when no trajectory is ready to use it. We introduce TreeRollout, a rollout engine that adapts trajectory generation to current GPU and tool bottlenecks by expanding policy-sampled trajectories into a tree. When tools are saturated, it branches after completed tool responses, reusing tool work to supply ready decoding work. When decoding becomes the bottleneck, the engine spawns new branches for tool calls from selected reasoning prefixes, enabling more frequent tool invocation with reduced reasoning overhead. We evaluate TreeRollout on different benchmarks covering math, search, and code generation. Given the same rollout time, TreeRollout yields significantly more trajectories than the baseline, with speedups of up to 1.82× in GRPO and 1.63× in RFT, and the largest gains appearing under severely imbalanced GPU and tool resources. Beyond raw throughput, these additional rollouts also lift the final trained model’s task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.