TRAINING RECURSIVE AGENT SWARMS AT DEPTH
Abstract
Agent swarms offer a promising path toward solving complex tasks through coordinated delegation and collective learning. However, most previous studies focus on agent coordination or shared policy optimization. How delegation depth and upstream learning jointly constrain training opportunities for independently updated subagents remains underexplored. In this study, we model this dependence through update coverage. This is the probability that a root rollout group supplies a policy with a usable update. We obtain two findings. Passive sampling costs can grow exponentially with delegation depth. In a finite batch model, individually beneficial upstream updates can jointly reduce downstream coverage and lower two-stage team value even under their best order. We introduce LACA (Learning Aware Controlled Adaptation) to explicitly elicit the training signals needed by nested policies. LACA samples fresh continuations from saved parent states and selects ordered policy subsets using update coverage and estimated two-stage team value. It retains the original task reward and learner objective. We evaluate LACA through paired update programs, controlled entry, and downstream policy rollback. At a common budget of 2,048 charged rollouts, LACA improves task success over GRPO by 8.42% on average across WebWalkerQA, Terminal-Bench 2.1, and BrowseComp-Plus, and its mean success exceeds PPO, DAPO, and Dr.MAS.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.