StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis

Axel SauerTero KarrasSamuli LaineAndreas GeigerTimo Aila

article2023ICML283 citations

Presents StyleGAN-T, a scalable GAN architecture that produces high-quality, text-aligned images in a single forward pass, outperforming distilled diffusion models in both generation speed and sample fidelity.

Listen

Recent advances in automated image generation from text prompts have relied predominantly on diffusion and autoregressive architectures. While these models deliver high visual quality, they require multiple iterative computing steps to produce a single image, making real-time deployment costly and slow. In contrast, generative adversarial networks—or GANs—generate images in a single forward pass, offering near-instantaneous output and flexible image editing, but have historically lagged behind in output quality and prompt alignment when applied to large, diverse datasets.

The article aims to evaluate whether a newly redesigned GAN architecture, named StyleGAN-T, can regain competitiveness and deliver state-of-the-art sample quality and strong prompt alignment in large-scale text-to-image synthesis at significantly higher operational speeds.

To achieve this, the authors redesigned the generator and discriminator components from an established GAN framework. They enhanced generator depth and conditioning mechanisms, simplified the discriminator using a vision transformer backbone, and introduced specialized training guidance using vision-language models alongside a latent truncation mechanism to balance image variation and prompt matching. The model was trained at scale on 250 million text-image pairs over four weeks using 64 high-performance graphics processors, representing approximately one-quarter of the compute budget used by competing diffusion models.

The findings establish that StyleGAN-T achieves dramatic improvements over prior GAN architectures and establishes a strong speed advantage over other frameworks. At a native resolution of 64×64 pixels, StyleGAN-T outperforms leading diffusion models on standardized quality benchmarks. When generating at 256×256 pixels, it operates at 10 frames per second (0.1 seconds per image), rendering images roughly 30 to 300 times faster than top diffusion baselines and outperforming distilled fast diffusion variants in both image fidelity and prompt alignment. However, at higher resolutions, its visual quality still trails behind top multi-step diffusion and autoregressive systems.

These results demonstrate that GAN-based text-to-image systems offer a viable, cost-effective alternative for high-throughput, low-latency, and interactive applications where current diffusion models are computationally prohibitive. Decision-makers should consider GAN architectures when application responsiveness, lower compute infrastructure costs, and smooth visual transitions between prompts are primary operational priorities.

For future work, the article recommends investigating enhanced high-resolution upscaling stages with expanded parameter capacity and longer training schedules to close the remaining high-resolution quality gap. Additionally, organizations should explore integrating larger language models to reduce occasional failures in text rendering and complex prompt attribute bindings.

arXiv: 2301.09515
Cover for StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis

Abstract

Text-to-image synthesis has recently seen significant progress thanks to large pretrained language models, large-scale training data, and the introduction of scalable model families such as diffusion and autoregressive models. However, the best-performing models require iterative evaluation to generate a single sample. In contrast, generative adversarial networks (GANs) only need a single forward pass. They are thus much faster, but they currently remain far behind the state-of-the-art in large-scale text-to-image synthesis. This paper aims to identify the necessary steps to regain competitiveness. Our proposed model, StyleGAN-T, addresses the specific requirements of large-scale text-to-image synthesis, such as large capacity, stable training on diverse datasets, strong text alignment, and controllable variation vs. text alignment tradeoff. StyleGAN-T significantly improves over previous GANs and outperforms distilled diffusion models - the previous state-of-the-art in fast text-to-image synthesis - in terms of sample quality and speed.

Table of Contents

  • 1 Introduction
  • 2 StyleGAN-XL
  • 3 StyleGAN-T
  • 3.1 Redesigning the Generator
  • 3.2 Redesigning the Discriminator
  • 3.3 Variation vs. Text Alignment Tradeoffs
  • 4 Experiments
  • 4.1 Quantitative Comparison to State-of-the-Art
  • 4.2 Evaluating Variation vs. Text Alignment
  • 4.3 Qualitative Results
  • 5 Limitations and Future Work
  • References
  • A Configuration Details
  • B Truncation Grids

Citation

MLA
Sauer, A., et al. “StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis”. arXiv, 2023, http://arxiv.org/abs/2301.09515v1.
APA
Sauer, A., Karras, T., Laine, S., Geiger, A., & Aila, T. (2023). StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis. arXiv. http://arxiv.org/abs/2301.09515v1
Chicago
Sauer, A., T. Karras, S. Laine, A. Geiger, and T. Aila. 2023. “StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis”. arXiv. http://arxiv.org/abs/2301.09515v1.
Harvard
Sauer, A. et al. (2023) “StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.09515v1.
Vancouver
1. Sauer A, Karras T, Laine S, Geiger A, Aila T (2023) StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis. arXiv

BibTeX

@article{sauer2023stylegan,
  title = {StyleGAN-T: Unlocking the Power of GANs for Fast Large-Scale Text-to-Image Synthesis},
  author = {Sauer, Axel and Karras, Tero and Laine, Samuli and Geiger, Andreas and Aila, Timo},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.09515v1},
  eprint = {2301.09515}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/