Dynamic Workflow Routing for Multi-Agent Systems
Abstract
We introduce a training-free framework for adapting workflows in large language model multi-agent systems at inference time. Rather than committing every request to a predetermined sequence of planning, generation, and refinement, this selects subsequent actions based on the outputs produced during execution. This flexibility addresses a limitation of fixed workflows: routine requests can incur unnecessary computation, while complex or ambiguous requests may receive inadequate reasoning and verification. Routing follows a two-tier design that balances decision cost with oversight. A lightweight Scout checks progress at each stage, escalating ambiguous or high-risk cases to a more capable LLM-based Router. These assessments determine whether the workflow should terminate, proceed, or introduce corrective steps, allowing computational effort to respond to the needs of each instance. Across the evaluated benchmarks, this achieves reductions of 46.1% in token consumption and 47.8% in end-to-end latency, demonstrating the efficiency benefits of adapting agent coordination during execution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.