VTD: Visible-Text Drafting for Speculative Decoding in Vision-Language Models
Abstract
Generating HTML or Markdown from images often requires vision-language models to produce long sequences, making autoregressive decoding a major source of latency. Retrieval drafters can accelerate generation without training, but their text sources may omit content visible in the image. At steps where dynamic retrieval accepts no draft descendant, only of the target's next three-token spans have a matching occurrence beginning at an earlier position in the same output. This limited recurrence motivates an additional retrieval source. We introduce VTD (Visible-Text Drafting), a speculative decoding method that uses OCR-recovered text as a retrieval source for drafting. We run OCR once, index the recognized text in a suffix automaton, and fuse its draft tree with a tree from a dynamic automaton over the prompt and accepted output. VTD requires no drafter training, external text corpus, or changes to target model weights. On five benchmarks, VTD achieves mean end-to-end speedups of on Qwen2.5-VL-3B, on Qwen2.5-VL-7B, and on InternVL2.5-8B. VTD outperforms every evaluated training-free retrieval baseline in speedup on every benchmark and target, and improves average speedup over the trained ViSpec drafter by on Qwen2.5-VL-7B. These results show that visible text is a complementary source for retrieval drafting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.