Co-Evolving Query Routing and Agent Specialization in Mixture-of-Agents
Abstract
Mixture-of-Agents systems typically treat query routing and agent post-training as separate stages. A router fitted to a frozen pool cannot track later policy updates, and agents are not given competence-aware data that would induce complementary expertise. We study this coupling as co-evolution: a router and a population of independently updated LLM agents are optimized in the same reinforcement learning loop, so that allocation and specialization change together. Before generation, a predictive familiarity estimator maps intermediate hidden states to a familiarity score that estimates how well an agent matches a query, without full rollouts or external judges. Cumulative-threshold routing then activates the smallest score-ranked subset whose familiarity scores sum to at least a threshold, trading accuracy against compute. The same scores steer training queries toward agents that are becoming proficient on related inputs, shifting a pool of homogeneous generalists into specialists. Across multiple domains, the resulting system outperforms static-agent routers and fine-tuning pipelines that keep the collaboration workflow fixed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.