keyword
generator architecture
A generator architecture is the structural design and configuration of neural network layers in a generative model that transforms input representations, such as random latent vectors or conditioning prompts, into synthetic data samples such as images, audio, or text. In generative adversarial networks and related generative frameworks, this component defines the specific arrangement of operations, including convolutions, upsampling, normalization, and attention mechanisms, that enable the model to synthesize high-dimensional outputs from lower-dimensional inputs. The design of a generator architecture directly influences training stability, computational efficiency, output fidelity, and the capacity to separately control high-level attributes and fine-grained variations in the generated media.
2 items

StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis
Axel Sauer, Tero Karras, Samuli Laine, Andreas Geiger, Timo Aila
Why you should read this
Presents StyleGAN-T, a scalable GAN architecture that produces high-quality, text-aligned images in a single forward pass, outperforming distilled diffusion models in both generation speed and sample fidelity.
Text-to-image synthesis has recently seen significant progress thanks to large pretrained language models, large-scale training data, and the introduction of scalable model families such as diffusion and autoregressive models. However, the best-performing models require iterative evaluation to generate a single sample. In contrast, generative adversarial networks (GANs) only need a single forward pass. They are thus much faster, but they currently remain far behind the state-of-the-art in large-scale text-to-image synthesis. This paper aims to identify the necessary steps to regain competitiveness. Our proposed model, StyleGAN-T, addresses the specific requirements of large-scale text-to-image synthesis, such as large capacity, stable training on diverse datasets, strong text alignment, and controllable variation vs. text alignment tradeoff. StyleGAN-T significantly improves over previous GANs and outperforms distilled diffusion models - the previous state-of-the-art in fast text-to-image synthesis - in terms of sample quality and speed.
Added
2026-09-28

A Style-Based Generator Architecture for Generative Adversarial Networks
Tero Karras, Samuli Laine, Timo Aila
Why you should read this
Redesigns the generator using style transfer principles (AdaIN) to separate high-level attributes from stochastic variation.
We propose an alternative generator architecture for generative adversarial networks, borrowing from style transfer literature. The new architecture leads to an automatically learned, unsupervised separation of high-level attributes (e.g., pose and identity when trained on human faces) and stochastic variation in the generated images (e.g., freckles, hair), and it enables intuitive, scale-specific control of the synthesis. The new generator improves the state-of-the-art in terms of traditional distribution quality metrics, leads to demonstrably better interpolation properties, and also better disentangles the latent factors of variation. To quantify interpolation quality and disentanglement, we propose two new, automated methods that are applicable to any generator architecture. Finally, we introduce a new, highly varied and high-quality dataset of human faces.
Added
2026-03-07
