MessageFlow: Proactive Message-Aware Execution for Efficient On-Device Multi-Agent Systems
Abstract
On-device multi-agent systems (MAS) powered by large language models (LLMs) face substantial response latency from chains of dependent LLM inference and tool calls. Agent-centric execution defers prompt processing until an agent is invoked, even when parts of its input are available earlier. We observe that evolving workflow messages can form stable, contiguous prompt prefixes that are ready for computation before control-flow dependencies are satisfied. Meanwhile, tool waits and autoregressive decoding often leave GPU resources underutilized, creating opportunities to overlap this early computation with ongoing execution. We present MessageFlow, a proactive message-aware execution framework that exploits these opportunities while preserving topology-defined agent invocation semantics. MessageFlow combines message update semantics, producer–consumer dependencies, and prompt ordering to identify reusable prefixes. Its runtime incrementally maintains and prioritizes the resulting prefill workloads, while an opportunity-aware GPU scheduler executes them during idle and decode-only periods and adapts to foreground demands to limit contention. The resulting key–value states are reused at invocation, moving part of prompt processing off the critical path. Implemented on LangGraph and SGLang, MessageFlow is evaluated on three representative workflows, four locally deployed LLMs, and three GPU platforms. Compared with state-of-the-art MAS execution systems, MessageFlow achieves up to 1.60× end-to-end speedup (1.32× on average), while increasing average GPU utilization from 57.8% to 89.6%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.