Accelerating Diffusion Transformers with Token-wise Feature Caching

Chang ZouXuyang LiuTing LiuSiteng HuangLinfeng Zhang

article2025ICLR154 citations

Proposes a training-free token-wise feature caching method that accelerates image and video diffusion transformers by up to 2.36x without sacrificing visual quality by adaptively caching features based on individual token sensitivity across network layers.

  • Paper: Scalable Diffusion Models with Transformers, William Peebles et al. (2023). Read this foundation for Diffusion Transformers first: it establishes the token-based architecture and scaling context that the source accelerates through feature caching.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Its formulation of iterative diffusion sampling clarifies the repeated denoising steps whose redundant transformer features the source caches.

No sufficiently relevant recommendations were found.

Abstract

Diffusion transformers have shown significant effectiveness in both image and video synthesis at the expense of huge computation costs. To address this problem, feature caching methods have been introduced to accelerate diffusion transformers by caching the features in previous timesteps and reusing them in the following timesteps. However, previous caching methods ignore that different tokens exhibit different sensitivities to feature caching, and feature caching on some tokens may lead to 10×\times more destruction to the overall generation quality compared with other tokens. In this paper, we introduce token-wise feature caching, allowing us to adaptively select the most suitable tokens for caching, and further enable us to apply different caching ratios to neural layers in different types and depths. Extensive experiments on PixArt-α\alpha, OpenSora, and DiT demonstrate our effectiveness in both image and video generation with no requirements for training. For instance, 2.36×\times and 1.93×\times acceleration are achieved on OpenSora and PixArt-α\alpha with almost no drop in generation quality.

Citation

MLA
Zou, C., et al. “Accelerating Diffusion Transformers with Token-wise Feature Caching”. arXiv, 2024, http://arxiv.org/abs/2410.05317v4.
APA
Zou, C., Liu, X., Liu, T., Huang, S., & Zhang, L. (2024). Accelerating Diffusion Transformers with Token-wise Feature Caching. arXiv. http://arxiv.org/abs/2410.05317v4
Chicago
Zou, C., X. Liu, T. Liu, S. Huang, and L. Zhang. 2024. “Accelerating Diffusion Transformers with Token-wise Feature Caching”. arXiv. http://arxiv.org/abs/2410.05317v4.
Harvard
Zou, C. et al. (2024) “Accelerating Diffusion Transformers with Token-wise Feature Caching”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2410.05317v4.
Vancouver
1. Zou C, Liu X, Liu T, Huang S, Zhang L (2024) Accelerating Diffusion Transformers with Token-wise Feature Caching. arXiv

BibTeX

@article{zou2024accelerating,
  title = {Accelerating Diffusion Transformers with Token-wise Feature Caching},
  author = {Zou, Chang and Liu, Xuyang and Liu, Ting and Huang, Siteng and Zhang, Linfeng},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2410.05317v4},
  eprint = {2410.05317}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors