acceptodds
Under review as a conference paper at ICLR 2027

AsyncSwarm: Asynchronous Multi-Agent Training and Inference for Information Seeking

Abstract

Applications of large language models (LLMs) have expanded from single-turn question answering to single-agent systems using multi-step reasoning and tools. Multi-agent systems further extend their ability to solve complex tasks. A classic multi-agent system design consists of a lead agent that assigns work to parallel sub-agents. We refer to this system as an LLM swarm, which is useful for complex information-seeking tasks that require extensive search to gather and verify evidence from multiple sources. However, blocking inference makes the lead agent wait for the slowest sub-agent. During online reinforcement learning (RL) for LLM swarms, rollouts also wait for sub-agents to finish dozens of tool calls. We propose AsyncSwarm with two components to address these challenges: Asynchronous Swarm Inference lets the lead agent launch new sub-agents as soon as any sub-agent finishes, while Trace-Retrieval Swarm RL retrieves stored sub-agent responses to replace slow online execution during training. We use the open-source GLM-4.5-Air and GLM-4.7-Flash as the lead and sub-agent models, respectively, and apply RL gradient updates only to the lead model. Compared with blocking baselines, AsyncSwarm achieves 1.6× faster inference with similar performance and 5.8× higher RL training throughput. Under the same RL time budget, AsyncSwarm improves over the initial swarm by 10.6 points on WideSearch and 7.6 points on BrowseComp-ZH, compared with gains of only 1.4 and 1.3 points from online RL, respectively. These results show that AsyncSwarm enables more efficient training while achieving substantially better performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.