acceptodds
Under review as a conference paper at ICLR 2027

EXACT WARM-STATE REUSE FOR HYBRID DIFFUSION LANGUAGE MODELS: CONTRACTS, SPEEDUPS, AND FAILURE BOUNDARIES

Abstract

Hybrid diffusion language models repeatedly revise a small set of tokens while most context is already stable, yet conventional inference can replay the entire sequence at every denoising step. Reuse is harder than causal prefix caching because a hybrid model combines paged attention state with recurrent Gated Delta Network (GDN) state, noncausal masks, absolute positions, and edit-dependent validity. We formulate a fail closed incremental execution contract over regions. A cache entry binds physical key-value pages and layer-local recurrent frontiers to token, position, model, adapter, mask-contract, and parent version identities. Our accepted GPU fast path is deliberately narrower than the general contract: a stable contiguous prefix and an active contiguous suffix, for which the attention/MLP active rows and the ordered GDN continuation are the same suffix. On one frozen HybridDiffusion- 2B checkpoint and one NVIDIA A30, warm reuse reduces measured serving-route CUDA latency from 1,055.27 to 137.39 ms (7.68×) for a 2,048-token prefix, 64-token suffix, and four diffusion steps, with bit-identical output hashes and 18.2% lower route-local peak memory. Across 36 controlled shapes, median warm speedup is 3.03×; 100 randomized region-contract cases pass their recorded correctness gates. We also report contrary evidence: natural generation is 11.99% slower, a KV-only route fails physical-page validation, and a four-region GPU route faults. We therefore claim an exact contiguous warm-state mechanism and a measured architecture envelope, not general fragmented execution, universal serving speedup, or cross-model transfer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.