AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning
Abstract
Online agentic reinforcement learning implemented with micro-services separates policy model training from rollout generation, improving scalability and modularity while potentially making frequent policy-weight synchronization a critical systems overhead. Shared storage can naturally connect these services across clusters, but vanilla dense policy-weight synchronization could be communication-intensive. Sparse synchronization reduces transferred data, yet checkpoint-oriented approaches can still retain a previous model and materialize complete intermediates to connect heterogeneous training and inference layouts. We present AReaL-DTE, a snapshot-free Delta Transfer Engine that turns inference-visible weight sparsity into end-to-end system efficiency. Across our evaluated workloads, fewer than 2% of BF16 weight elements change between consecutive policy versions. AReaL-DTE reconstructs overwritten weights on demand by inverting AdamW updates, streams reconstructed and current parameters through converter-aligned BF16 change detection, and remaps changed elements directly into receiver-local coordinates. AReaL-DTE supports manifest-committed sparse transfer through shared storage across clusters and a deadlock-safe two-round protocol within a cluster, followed by direct application to inference shards. We evaluate AReaL-DTE on Qwen3-8B and Qwen3-30B-A3B across four online RL workloads. AReaL-DTE achieves speedups of up to over ByteCheckpoint and over PULSE across clusters, and up to and , respectively, within a cluster. In the same-cluster Qwen3-30B-A3B experiments, it reduces peak GPU memory by approximately 41% and peak CPU memory by at least 87%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.