acceptodds
Under review as a conference paper at ICLR 2027

Speculative Multi-Trajectory Decoding for Efficient Multimodal LLM Planning

Abstract

Multimodal large language model (MLLM) planners typically serialize trajectories into token sequences and decode them autoregressively, making sequential generation a major inference bottleneck that becomes more severe when multiple candidates are required. To address this issue, we propose Speculative Multi-Trajectory Decoding (SMTD), a target-verified framework that accelerates trajectory generation through a unified draft–verify–select process while keeping the pretrained OneVL target model frozen. Lightweight Medusa and anchored-GRU drafters first propose blocks of future trajectory tokens, which are verified in parallel by the target model to reduce sequential evaluations while preserving target-consistent generation. Building on this mechanism, SMTD extends speculative decoding to multi-trajectory planning via speculative multi-sampling, efficiently producing multiple verified candidates. A Unified Online Candidate Selector (UOCS) then scores these candidates and selects the final trajectory at inference time. Experiments on NAVSIM, ROADWork, and Impromptu demonstrate consistent latency reductions with competitive planning quality. On NAVSIM, SMTD reduces mean decoding latency from 5.221 s to 1.349 s. The remaining gap between selected and oracle candidates further identifies online selection as a key bottleneck after generation is accelerated.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.