Accelerating Agentic Multimodal LLMs via Asynchronous Tool-Trajectory Speculation
Abstract
Agentic multimodal large language models use tools to acquire additional visual evidence, but repeated alternation between large-model inference and tool execution accumulates latency. Replacing the entire agent with a smaller model can reduce accuracy, motivating a division of work between exploratory action generation and final answering. We introduce , an asynchronous speculative execution framework in which a small draft model generates and executes tool actions while a large main model assesses the resulting evidence and remains responsible for answering. Separate quality and answerability assessments guide continued exploration, early answering, and main-model takeover. Assessment overlaps with draft exploration, while takeover reuses completed tool interactions to limit repeated work. Across five benchmarks, SpecAgent achieves – end-to-end speedup over the Qwen3.5-27B agent while maintaining broadly comparable accuracy. These results demonstrate an accuracy–latency trade-off through small–large model collaboration in tool-assisted multimodal reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.