Compass: Repairing When Target-Model Guidance Hurts in Parallel Speculative Decoding
Abstract
Target-model features are increasingly used to improve parallel speculative drafting, but stronger guidance does not necessarily translate into faster decoding. We reveal that target-model guidance can hurt it in two ways: feature injection adds drafting overhead that can limit the speed benefits of improved acceptance, while target-guided adaptation can introduce new errors in prefixes already accepted by the original drafter, undermining the benefit of target-model guidance. We introduce *Compass*, a lightweight target-guided parallel speculative decoding framework that combines compact layerwise target guidance with reference-aware guidance optimization. The guidance module compresses multi-layer target states into a shared low-dimensional source and injects separate gate and up input shifts into each draft MLP, reusing native projections to limit drafting overhead. The optimization module uses the initial drafter as a frozen reference model to protect previously accepted prefixes and encourage recovery beyond its acceptance boundary. We construct a margin-based prefix objective aligned with greedy verification and derive a per-block lower bound linking the designed loss to acceptance gain over the reference model. Implemented in vLLM, *Compass* achieves mean AR-relative speedups of 5.27, 4.25, and 2.39 on Qwen3-8B, Qwen3-14B, and Qwen3-32B across six benchmarks, exceeding the strongest available baseline for each target by 7.0%, 8.1%, and 12.4%. Ablations show that compact guidance retains acceptance benefits at lower drafting cost, while reference-aware optimization extends accepted prefixes and accelerates decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.