LiteCa: Memory-Efficient Feature Caching for Diffusion Acceleration
Abstract
Feature caching accelerates Diffusion Transformers (DiTs) by reusing intermedi- ate features across timesteps. Recently Forecasting-based methods show that mod- ule features evolve predictably across timesteps, enabling temporal extrapolation for feature forecasting. However, storing module-level features and their histori- cal states introduces substantial cache memory overhead. To solve this problem, we propose LiteCa, a training-free and plug-and-play memory-efficient feature caching framework. We observe that the Adaptive Layer Normalization (AdaLN) modulation parameters exhibit a low-rank structure, enabling LiteCa to build a more compact cache representation with an adjustable rank for different memory budgets. We further introduce Test-Time Correction to improve generation qual- ity without changing the original forecasting method. Experiments on recent DiT models show that LiteCa reduces feature-cache memory by 90–99% compared with module-level caching, while improving both inference speed and generation quality and achieving state-of-the-art performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.