RELOAD: Packed Future-State Re-Encoding for Low-Latency Exact Speculative Diffusion
Abstract
Exact speculative diffusion preserves a designated discretized target distribution, but autoregressive drafting leaves sequential neural calls on the critical path. Predicting a whole horizon from one anchor removes this rollout yet leaves later predictions without representations of their random future source states. We call this mismatch future-state representation staleness. RELOAD, a two-pass drafter, first constructs causal future sources from parallel coarse predictions. A packed second pass re-encodes these sources, and the scheduler rebuilds proposals from the same anchor and innovations. Rank-eight Role-LoRA specializes this second role with 267,264 trainable parameters. We prove that the resulting proposals retain prefix causality, allowing the unchanged reflection verifier to preserve the target-chain law in exact arithmetic under matched-covariance assumptions. At horizon eight, re-encoding raises mean accepted length from 2.02 to 4.52, and role adaptation further raises it to 4.77 in a one-GPU component study. On 256 paired ImageNet-256 cases, RELOAD reduces mean sampling latency by 10.0% relative to reproduced strict FREE, a 1.111× speedup under the same eight-rank allocation. Eight-way verification yields a 2.286× latency speedup over one-active-rank TARGET.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.