acceptodds
Under review as a conference paper at ICLR 2027

ProSpect: Proactive Speculation for Lossless Acceleration of Multi-Agent Systems

Abstract

Multi-agent LLM systems are increasingly used for complex, long-horizon tasks, but their sequential, dependent calls create high end-to-end latency. Speculative execution reduces latency losslessly, by pre-executing likely future calls and committing only those the system would have made anyway. Extending it from a single agent to a multi-agent system is hard because a speculation must predict which agents are invoked next in an execution trace that is a DAG instead of a chain, and what each predicted agent receives. An agent starts early only if every incoming edge is predicted correctly. We present ProSpect, which separates these two predictions. Choosing the next agent is cast as a link prediction task over a graph of agents assembled from past traces and answered by a recurrence folded along the execution DAG; the Speculator LLM is left with the task of message generation alone for the predicted agent, and boosted further through retrieval-augmented in-context learning on Actor's past traces. The speculation budget between these two prediction heads is split between them adaptively, by the predictor's confidence. Across four multi-agent systems and spanning the Qwen, Gemma and DeepSeek LLM families, ProSpect saves up to of wall-clock time and outperforms the state of the art speculative acceleration method on every system, while also reducing token cost and compute flops.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.