Accelerating Diffusion Transformers with Token-wise Feature Caching
Chang ZouXuyang LiuTing LiuSiteng HuangLinfeng Zhang
Proposes a training-free token-wise feature caching method that accelerates image and video diffusion transformers by up to 2.36x without sacrificing visual quality by adaptively caching features based on individual token sensitivity across network layers.
- Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). Read this foundation for Diffusion Transformers first: it establishes the token-based architecture and scaling context that the source accelerates through feature caching.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Its formulation of iterative diffusion sampling clarifies the repeated denoising steps whose redundant transformer features the source caches.
No sufficiently relevant recommendations were found.