HASTE: Hierarchical Adaptive Spatio-Temporal Reuse for Efficient Image Decomposition
Abstract
Text-to-layer (T2L) diffusion models achieve editable image decomposition by jointly denoising multi-layer RGBA stacks, but suffer from prohibitive inference latency. Existing diffusion accelerators overlook this multi-layer structure, treating images as flat, spatially contiguous domains. We observe that computational redundancy in T2L models is inherently hierarchical, governed by inter-layer temporal heterogeneity (layers stabilize at disparate denoising stages) and intra-layer spatial sparsity (active synthesis is strictly confined to a sparse alpha support, outside of which latent states quickly freeze). To exploit this insight, we propose HASTE, the first dedicated online caching framework for T2L generation. HASTE infers alpha support online via an ultra-lightweight token-level regressor and deploys a two-stage hierarchical scheduler: converged layers bypass Transformer execution entirely via flow-matching endpoint extrapolation, while active layers restrict fresh evaluation strictly to tokens within the alpha support. To prevent cross-layer context fragmentation, HASTE executes a sparse DiT scheme powered by persistent Key-Value caches. On the Crello benchmark, HASTE accelerates Qwen-Image-Layered sampling by with negligible quality loss. Furthermore, it generalizes training-free across four diverse T2L backbones, delivering consistent – speedups. Code is available at https://anonymous.4open.science/r/HASTE-0434.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.