VFlash-FM: Training for One-Pass Drafting in Video Language Models
Abstract
Speculative decoding reduces the serial cost of video language model generation, but drafting itself is constrained by the long visual context: autoregressive drafters pay for it at every proposed token, while one-pass parallel drafters lose most of a block to an early error. We introduce VFlash-FM, which improves one-pass drafting through training alone. Our key observation is that a discrete-flow path from a fully masked block to the target continuation defines a family of auxiliary prediction tasks: completing a block from the drafter's own partially revealed predictions, including its mistakes. Training the drafter on these self-generated partial states, with a step-aware flow objective and target-confidence weighting oriented toward accepted prefixes, improves the same network's predictions from all-mask inputs. The extra computation is paid once during training rather than at every decoding step, so deployment remains a single drafter forward over cached target features followed by target verification. Across five video benchmarks, VFlash-FM reaches macro decoding speedups of on LLaVA-OneVision-7B and on Qwen2.5-VL-7B, exceeding the published speedups of prior video drafting methods we compare against; on LLaVA, it improves over a non-flow drafting variant in both committed progress and speed across all tested verification configurations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.