Built independently by an author, for readers. Read the story and support ChapterPal

keyword

representation alignment

Representation alignment is a machine learning technique where the internal feature embeddings or hidden states produced by a model are trained to correspond with or match target representations from another network, modality, or reference feature space. By constraining intermediate layers to align with semantically rich representations, such as those derived from robust pretrained visual or linguistic encoders, the target model can leverage existing structured knowledge rather than learning complex feature abstractions from scratch. This regularization mechanism is widely applied across multimodal learning, domain adaptation, and generative modeling to accelerate training convergence, improve semantic consistency, and enhance the overall quality and stability of model outputs.

1 item

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, Saining Xie

OrganizationsKorea Advanced Institute of Science and TechnologyKorea UniversityNew York UniversityScaled Foundations

Why you should read this

Introduces REPA, a simple regularization framework that achieves over 17.5x faster training convergence and state-of-the-art image generation quality by aligning diffusion transformer representations with those from pretrained self-supervised visual encoders.

Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that one main bottleneck in training large-scale diffusion models for generation lies in effectively learning these representations. Moreover, training can be made easier by incorporating high-quality external visual representations, rather than relying solely on the diffusion models to learn them independently. We study this by introducing a straightforward regularization called REPresentation Alignment (REPA), which aligns the projections of noisy input hidden states in denoising networks with clean image representations obtained from external, pretrained visual encoders. The results are striking: our simple strategy yields significant improvements in both training efficiency and generation quality when applied to popular diffusion and flow-based transformers, such as DiTs and SiTs. For instance, our method can speed up SiT training by over 17.5×\times, matching the performance (without classifier-free guidance) of a SiT-XL model trained for 7M steps in less than 400K steps. In terms of final generation quality, our approach achieves state-of-the-art results of FID=1.42 using classifier-free guidance with the guidance interval.

Added

2026-05-18