PRISM: PRIvileged Sibling-guided Multi-Agent Topology Design
Abstract
The effectiveness of large language model multi-agent systems depends on how agents are selected and connected for each query. Autoregressive topology designers learn these decisions from execution-verified query-graph pairs, but continued supervised training does not necessarily improve downstream performance. In our experiments, continued supervised fine-tuning increases output concentration, with uneven performance across benchmarks. We propose PRISM, which reuses the verified corpus through on-policy self-distillation from a privileged teacher. During training, the teacher receives a different graph that successfully solved the same query through a low-rank conditioning branch. The student samples topologies under its own policy and learns from the teacher's distributions on these trajectories, together with a supervised anchor on verified graphs. The teacher is discarded at deployment, leaving the designer architecture unchanged, and post-training requires no additional LLM executions. Across six reasoning and code benchmarks with Qwen2.5-7B-Instruct, PRISM outperforms the evaluated topology-design baselines. Matched teacher controls support the contribution of query-associated privileged information, while output-distribution measurements show reduced concentration relative to continued supervised fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.