keyword
diffusion transformers
Diffusion transformers are a class of generative artificial intelligence models that replace the conventional convolutional U-Net architecture in diffusion models with a transformer-based neural network backbone. In these systems, data such as images, videos, or their compressed latent representations are divided into sequences of flattened patches and processed as tokens through self-attention mechanisms. The model is trained to iteratively predict and remove noise from corrupted inputs to synthesize new, high-fidelity data, integrating conditioning signals like text prompts, class labels, and diffusion timesteps via specialized modulation or attention layers. By applying the transformer architecture to generative diffusion processes, diffusion transformers exhibit strong scaling properties, where increases in model parameters, token counts, and training compute systematically yield improvements in generation quality and computational efficiency.
3 items

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, Saining Xie
Why you should read this
Introduces REPA, a simple regularization framework that achieves over 17.5x faster training convergence and state-of-the-art image generation quality by aligning diffusion transformer representations with those from pretrained self-supervised visual encoders.
Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that one main bottleneck in training large-scale diffusion models for generation lies in effectively learning these representations. Moreover, training can be made easier by incorporating high-quality external visual representations, rather than relying solely on the diffusion models to learn them independently. We study this by introducing a straightforward regularization called REPresentation Alignment (REPA), which aligns the projections of noisy input hidden states in denoising networks with clean image representations obtained from external, pretrained visual encoders. The results are striking: our simple strategy yields significant improvements in both training efficiency and generation quality when applied to popular diffusion and flow-based transformers, such as DiTs and SiTs. For instance, our method can speed up SiT training by over 17.5, matching the performance (without classifier-free guidance) of a SiT-XL model trained for 7M steps in less than 400K steps. In terms of final generation quality, our approach achieves state-of-the-art results of FID=1.42 using classifier-free guidance with the guidance interval.
Added
2026-05-18

Masked Autoencoders Are Effective Tokenizers for Diffusion Models
Hao Chen, Yujin Han, Fangyi Chen, Xiang Li, Yidong Wang, Jindong Wang, Ze Wang, Zicheng Liu, Difan Zou, Bhiksha Raj
Why you should read this
Presents MAETok, an autoencoder that significantly boosts diffusion model efficiency and generation quality by learning semantically rich latent spaces, achieving state-of-the-art ImageNet generation with 76x faster training and 31x higher inference throughput.
Recent advances in latent diffusion models have demonstrated their effectiveness for high-resolution image synthesis. However, the properties of the latent space from tokenizer for better learning and generation of diffusion models remain under-explored. Theoretically and empirically, we find that improved generation quality is closely tied to the latent distributions with better structure, such as the ones with fewer Gaussian Mixture modes and more discriminative features. Motivated by these insights, we propose MAETok, an autoencoder (AE) leveraging mask modeling to learn semantically rich latent space while maintaining reconstruction fidelity. Extensive experiments validate our analysis, demonstrating that the variational form of autoencoders is not necessary, and a discriminative latent space from AE alone enables state-of-the-art performance on ImageNet generation using only 128 tokens. MAETok achieves significant practical improvements, enabling a gFID of 1.69 with 76x faster training and 31x higher inference throughput for 512x512 generation. Our findings show that the structure of the latent space, rather than variational constraints, is crucial for effective diffusion models. Code and trained models are released.
Added
2026-04-19


Scalable Diffusion Models with Transformers
William Peebles, Saining Xie
Why you should read this
Introduces Diffusion Transformers (DiTs), a new class of diffusion models that leverage transformer architectures, demonstrating state-of-the-art image generation performance and superior scalability, fundamentally challenging the dominance of U-Net backbones in diffusion models.
We explore a new class of diffusion models based on the transformer architecture. We train latent diffusion models of images, replacing the commonly-used U-Net backbone with a transformer that operates on latent patches. We analyze the scalability of our Diffusion Transformers (DiTs) through the lens of forward pass complexity as measured by Gflops. We find that DiTs with higher Gflops -- through increased transformer depth/width or increased number of input tokens -- consistently have lower FID. In addition to possessing good scalability properties, our largest DiT-XL/2 models outperform all prior diffusion models on the class-conditional ImageNet 512x512 and 256x256 benchmarks, achieving a state-of-the-art FID of 2.27 on the latter.
Added
2026-02-25

