Beyond Monolithic Architectures: Multi-Agent Search and Knowledge Optimization for Agentic Search
Abstract
Agentic search has emerged as a promising paradigm for complex information seeking by enabling Large Language Models (LLMs) to interleave reasoning with tool use. However, prevailing systems largely rely on monolithic agents, which suffer from structural bottlenecks such as unconstrained reasoning outputs that inflate trajectories, sparse outcome-level rewards that hinder credit assignment, and stochastic search noise that destabilize learning. To address these challenges, we propose M-ASK (Multi-Agent Search and Knowledge), a framework for multi-agent joint optimization in agentic search. M-ASK decomposes the search process into two complementary groups of trainable agents: Search Behavior Agents, which govern planning and search actions, and Knowledge Management Agents, which refine, compress, and maintain the evolving internal context. Crucially, this role decomposition is not introduced merely for modular orchestration; instead, M-ASK jointly optimizes these heterogeneous agents within a unified training framework, so that search behavior and knowledge management are learned as a coordinated multi-agent policy. To make such coordination effective, M-ASK employs turn-level rewards that provide fine-grained supervision for intermediate search decisions and knowledge updates. Experiments on multiple QA benchmarks demonstrate that M-ASK consistently outperforms strong baselines, yielding both higher answer accuracy and substantially improved training stability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.