acceptodds
Under review as a conference paper at ICLR 2027

From Delegation to Coordination: Learning Efficient Multi-Agent Deep Research

Abstract

Multi-agent deep-research systems use delegation to parallelize search and keep long tool traces out of the main agent’s context. Yet delegation helps only when task allocation and subtask execution are coordinated: the main agent must issue useful subtasks, while subagents must return reliable evidence. Existing train- ing either optimizes only the main agent or broadcasts the same team advantage to all agents, which can reward misleading subagent reports and penalize use- ful ones. We introduce LEAD-RL (Learning Effective Agent Delegation with Reinforcement Learning), a joint main-agent–subagent training framework with contribution-aware credit routing. A training-time rubric evaluates the research utility of each subagent report and gates its inherited team advantage, while final- answer correctness remains the only team reward; no token, step, or speedup re- ward is used. Across five benchmarks and two model families, LEAD-RL im- proves answer quality. On BrowseComp, it improves accuracy by 2.49 percentage points and reduces critical-path cost by 24.5% over joint training with shared team credit. During validation, accuracy increases by 6.78 points at approximately con- stant critical-path cost. These results show that contribution-aware credit assign- ment improves the quality–efficiency trade-off of multi-agent deep research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.