Do Multi-Agent LLMs Explore Each Other?
Abstract
Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can explore effectively when interacting with one another. We show that modern LLM agents often fail to do so, exhibiting myopic and polarized interaction patterns that lead to suboptimal coordination and increased regret. We formalize this challenge as the Multi-Agent Exploration problem, modeling it as a partially observable stochastic game (POSG), in which agents must probe peers to infer their capabilities and identify effective interaction strategies. To address this, we introduce Multi-Agent Contextual Exploration (MACE), a lightweight framework that explicitly promotes exploration through structured peer selection. Across both contextual and parametric diversity settings, exploration behavior and downstream task performance substantially improve when MACE is applied. We further provide theoretical insight that MACE achieves a sublinear expected-regret bound. Overall, our results highlight a fundamental limitation of current LLM agents and underscore the importance of explicitly guided exploration for reliable multi-agent autonomy. Code will be released for reproducible research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.