acceptodds
Under review as a conference paper at ICLR 2027

CSRe: Training-Free Diffusion Transformer Acceleration via Component-Wise Selective Refresh

Abstract

Diffusion Transformers (DiTs) repeatedly evaluate attention and MLP modules across denoising steps, despite temporal redundancy in their internal features. Feature caching reduces this cost, but fine-grained reuse must account for both refresh allocation and execution overhead. We introduce CSRe (Component-Wise Selective Refresh), a training-free inference framework that selectively recomputes output-projected attention-head and MLP hidden-channel contributions. Full anchor steps refresh the caches and measure changes between anchors; these lagged measurements guide component selection for subsequent selective steps. Selected contributions update the cached module outputs through their latest-cache deltas. A shared compute budget coordinates attention and MLP refresh, with the remaining MLP budget distributed across blocks according to anchor-measured hidden-activation changes. Selected channels are executed using hardware-aligned dense submatrices. We evaluate CSRe on DiT-XL/2, PixArt-α, and FLUX.1-dev. On an RTX 5090D, CSRe maintains competitive generation quality while achieving denoiser speedups of 2.292×, 1.963×, and 1.728× on DiT-XL/2, PixArt-α, and FLUX.1-dev, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.