PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting
Abstract
In a multi-agent LLM system, the main agent must write each sub-agent's prompt, yet current models struggle to determine what each sub-agent needs to know. Existing benchmarks score belief tracking or downstream task success, not the prompts themselves. We introduce PerspectiveGap, a benchmark for evaluating LLMs' ability to compose orchestration prompts for multi-agent systems. PerspectiveGap contains 110 scenarios, each evaluated through two distractor-mixed task formats: role-fragment assignment and verbatim prompt assembly, in which the model assembles each role's prompt from the given fragments. These scenarios are organized into 10 loop-centered topologies distilled from the authors' real-world engineering practice. Across 35 commercial models from 10 companies, the average combined pass rate is only 17.6%, and in role-fragment assignment, models give roles an average of 2.1 out-of-role fragments per scenario. GPT-5.5 leads by a wide margin (62.0% pass rate). These results show that current models often fail to give each sub-agent exactly the context it needs, an ability that existing benchmarks do not measure.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.