Centroid-Preserving Dynamics: Composable Asynchronous Coordination for Multi-Learner Reinforcement Learning
Abstract
While multi-learner parallelism scales training throughput in deep reinforcement learning (DRL), asynchronous coordination often fails to translate raw throughput into policy improvements. We show that coordination mechanisms can shift the population parameter centroid and induce structural drift, destabilizing the coupled parameter-policy-data feedback loop. To address this, we model multi-learner training as a driven-dissipative dynamical system. Through an orthogonal decomposition into parameter centroid and consensus disagreement, we identify two essential operator-theoretic criteria for stable coordination: centroid preservation and disagreement non-expansion. We prove both properties are closed under finite asynchronous compositions, establishing asymptotic disagreement bounds under window contraction and closed-loop feedback. Guided by these principles, we develop a centralized 1/N-normalized incremental commit mechanism and a decentralized symmetric zero-sum pairwise protocol, accompanied by an in-flight mass model capturing physical delays. Across the TianJi benchmark and four Atari environments, scaling to three nodes demonstrates that our decentralized protocol achieves a 1.66×–2.05× wall-clock speedup over synchronous baselines with higher tail returns and AUC. Furthermore, it cuts time-to-threshold by 10.9%–46.1% over non-centroid-preserving baselines and maintains robust stability under dynamic network latencies, showing that principled coordination reliably converts barrier-free throughput into end-to-end efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.