SQB-Spec: Accelerating Large Multimodal Models through Task-Specific Speculative Decoding
Abstract
The deployment of Vision-Language Models (VLMs) is severely limited by the high latency of autoregressive decoding. While speculative decoding offers a lossless acceleration solution, existing monolithic draft models encounter a "representation ceiling" when handling the inherent task heterogeneity of VLMs. In this paper, we propose SQB-Spec, a task-specific speculative decoding framework for efficient and robust VLM acceleration. Built upon the EAGLE-3 paradigm, SQB-Spec introduces modality-specific projection pathways for enhanced cross-modal alignment and incorporates a Gated Attention mechanism to mitigate "attention sinks" in long-sequence multimodal drafting. To address task heterogeneity, we further propose the Specialized Query-projection and Bias (SQB) mechanism, which modularizes task-specific representational power into lightweight Query-projection branches and linear biases. A phase-level router adaptively activates these pathways, guided by a dual-component objective: a Historical Load Balancing Loss for global expert utilization and a Stride-based Temporal Smoothness Loss for aligning token-level training with stationary inference. Extensive evaluations on LLaVA-OneVision and Qwen3-VL series demonstrate that SQB-Spec achieves a peak speedup of 2.60x while strictly maintaining the original output distribution, significantly outperforming state-of-the-art baselines across diverse benchmarks including MMMU, DocVQA, and MVBench.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.