acceptodds
Under review as a conference paper at ICLR 2027

SparseDFlash: Dynamic Context Retrieval for Long-video Speculative Decoding

Abstract

Block-diffusion drafters accelerate speculative decoding through parallel token prediction, but repeatedly attending to long video contexts can make drafting a bottleneck to overall speedup. Existing approaches reduce this overhead by re- stricting the draft context through sliding-window attention or visual token prun- ing. However, permanently discarding context can remove evidence needed for subsequent predictions. This limitation is particularly relevant to long videos, where successive predictions may depend on different, temporally distant events. We introduce SparseDFlash, an attention framework for efficient long-video spec- ulative decoding that reuses the drafter’s existing attention parameters without in- troducing additional learnable parameters. The framework preserves the complete context while dynamically restricting attention to evidence relevant to each draft- ing step. This allows the drafter to revisit distant evidence as generation progresses without repeatedly attending to the entire context. We further address a limitation of naive raster-order block construction, which can separate spatiotemporally ad- jacent visual tokens across retrieval blocks. Locality-aware ordering based on a three-dimensional Hilbert curve organizes these blocks according to the video’s spatial and temporal structure. The framework preserves the target model’s output distribution through standard speculative verification. Experiments on long-video description tasks using videos from four video understanding benchmarks show that SparseDFlash achieves state-of-the-art performance across all four bench- marks with both Qwen3-VL-8B and Qwen3.5-9B, while consistently outperform- ing the full-attention drafter using the same checkpoint.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.