Drafting Without Hearing: Why Speculative Decoding Fails in Hierarchical Speech Generation
Abstract
Speech models such as Qwen3-Omni represent each audio frame with a primary code followed by residual codes that refine its acoustic detail. The completed frame becomes part of the history used to predict the next frame. Speculative decoding offers a natural way to accelerate this process by verifying future proposals in parallel. On frozen Qwen3-Omni, scoring three positions costs only 23% more than scoring one, yet complete speculative generation is consistently slower than ordinary decoding. We trace the lost verification saving to generating each draft's remaining acoustic codes, a step we call acoustic completion. Those codes are needed to verify later proposals before the draft's acceptance is known. Preparing these inputs and completing output frames account for 81.7% of measured speculative model computation. Even with a free drafter, the remaining computation nearly matches ordinary decoding at the observed acceptance rate. Truncating acoustic feedback to reduce this cost distorts subsequent target predictions. A controlled intervention raises word error rate from 11.68% to 40.97% even though every output frame retains all its codes. Drafting complete acoustic groups transfers the same prediction task to the drafter. On Moshi, more accurate complete-group drafters increase the number of accepted code groups, yet their candidate-generation cost still prevents a speedup. These results identify acoustic completion as a cost that higher proposal accuracy alone does not eliminate: a draft may be cheap to verify, yet expensive to make verifiable. Under the conditions we study, ordinary decoding is the better policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.