REPRISE: Efficient Long-Context Diffusion Inference via Distillation through Reused KV Caches
Abstract
Bidirectional diffusion language models repeatedly refresh key–value (KV) caches over long prompts during generation, incurring substantial computational cost. Sparse attention and cache reuse reduce this cost but alter the context states used for prediction, which can reduce generation quality. To address these limitations, we propose REPRISE, which adapts models to sparse attention and cache reuse through dense-model distillation and selective replay. During training, the student makes predictions using sparse attention and reused caches. A frozen dense teacher supervises these predictions based on the student’s current token sequence. Selective replay rebuilds the required cache computations so that later prediction losses can train cache construction without retaining the full generation graph. During inference, REPRISE refreshes prompt states less often than generation states and uses sparse attention for prompt refreshes. Compared with the widely used Prefix Cache and Dual Cache, REPRISE improves average scores by 8.72 and 7.67 points, respectively, while reducing time to first token (TTFT) by 48% and time per output token (TPOT) by 80–82% on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.