SpecKV: Lossless Speculative Decoding via Dynamic KV Selection for Efficient Inference of Large Vision-Language Models
Abstract
Large vision-language models (LVLMs) support increasingly complex multimodal tasks, but long visual inputs and growing textual histories make autoregressive decoding costly. While existing approaches reduce drafting costs through visual context compression, one-shot selection overlooks changing context requirements and information overlap across modalities. We analyze two characteristics of multimodal drafting: inter-step context heterogeneity and intra-step context redundancy. Our analysis reveals that context requirements shift across drafting rounds, while visual evidence can overlap with information already expressed in the textual context, limiting the effectiveness of fixed visual subsets. To address these challenges, we propose SpecKV, a training-free multimodal speculative decoding framework that dynamically selects visual and textual key–value (KV) entries under a shared drafting budget. SpecKV combines verification-guided context adaptation with text-conditioned visual complementarity, using the latest accepted target state to assess relevance and a relevance-weighted textual reference to estimate cross-modal overlap. The active context is refreshed after each verification round, while inactive entries remain available for reselection. Experiments on seven image and video understanding benchmarks across Qwen2.5-VL and LLaVA-OneVision demonstrate acceleration in both standard and self-speculative settings. With a draft KV budget of , SpecKV achieves up to decoding speedup on image tasks and on video tasks over vanilla autoregressive decoding, while preserving the target output distribution through exact verification with full context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.