FRAME: Failure-Aware Multi-Agent Topology Learning with Random Fourier Features
Abstract
Large language model (LLM)-based multi-agent systems (MAS) increasingly rely on learned collaboration topologies to coordinate specialized agents efficiently. Yet, existing topology optimizers remain largely task-conditioned: they construct topology from the task and predefined role capabilities, while overlooking evidence revealed during execution and providing limited mechanisms to recover when the resulting collaboration fails. To mitigate this drawbacks, we introduce FRAME (Failure-Aware Multi-Agent Topology Learning with Random Fourier Features), a failure-aware, execution-conditioned framework that unifies topology construction and repair through bilevel optimization. At the upper level, FRAME learns the collaboration graph autoregressively, conditioning each decision on the task, selected roles, and evolving execution states. Specifically, FRAME introduces Random Fourier Feature Upper Confidence Bound (RFF-UCB) role selection, combining nonlinear utility estimation with uncertainty-aware exploration, while communication edges are learned using a lightweight neural network optimized directly from task reward. When performance falls below a prescribed threshold, the lower level localizes the source of underperformance and performs targeted role repair, which is subsequently fed back into topology construction. This design turns topology learning from a one-shot graph prediction problem into a closed-loop process of constructing, executing, diagnosing, and repairing. Across diverse benchmarks, FRAME consistently outperforms the existing topology design algorithms while being token efficient.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.