StarGAN v2: Diverse Image Synthesis for Multiple Domains
Yunjey ChoiYoungjung UhJaejun YooJung-Woo Ha
Proposes a scalable image-to-image translation framework that generates diverse images across multiple visual domains within a single model and introduces the AFHQ benchmark dataset for evaluating cross-domain synthesis.
Modern computer vision applications increasingly rely on image-to-image translation to adapt visual content across distinct categories, such as changing human facial attributes or animal species. However, existing translation frameworks face a critical trade-off: methods capable of generating diverse styles typically only work between two categories and require training separate models for every category pair, while scalable multi-domain models tend to produce fixed, uniform outputs that lack style diversity. Addressing both scalability and output variety within a single system is essential for efficient and flexible generative image synthesis.
The article demonstrates that a unified generative framework, named StarGAN v2, can simultaneously achieve high visual quality, extensive style diversity, and multi-domain scalability using a single generator. It sets out to evaluate this approach across diverse image domains and benchmark its performance against leading multi-modal translation models.
The authors developed an architecture that replaces rigid domain labels with domain-specific style codes extracted from reference images or generated from random inputs. To evaluate this approach, the team conducted rigorous quantitative benchmarks and human perceptual studies using the human face dataset CelebA-HQ and a newly introduced 15,000-image Animal Faces-HQ dataset spanning cats, dogs, and wildlife. Output quality was measured via the standard Fréchet inception distance metric, while image diversity was assessed through learned perceptual image patch similarity and blind human preference surveys on Amazon Mechanical Turk.
The evaluation yielded several key findings. First, StarGAN v2 substantially improved image fidelity, achieving Fréchet inception distance scores of 13.7 on CelebA-HQ and 16.2 on Animal Faces-HQ in random-generation mode, outperforming the leading baseline by more than twofold. Second, in reference-guided translation, the model reduced error distances by approximately 1.5 times on human faces and 3.5 times on animal faces compared to the next best method. Third, the framework preserved source characteristics, such as pose and identity, while transferring fine-grained target styles like hairstyles, breeds, and eye colors. Finally, human evaluators decisively preferred StarGAN v2 over competing methods, awarding it between 68.9% and 92.1% of votes across visual quality and style-transfer categories.
These findings indicate that generative image systems do not need to trade architectural simplicity for stylistic variety. By relying on a single network for multiple categories, engineering teams can significantly lower deployment overhead, operational complexity, and the computational costs associated with training numerous pairwise models. Furthermore, the architecture successfully generalizes to unseen source distributions, making it a robust foundation for practical creative tools and automated visual content generation.
Organizations developing generative visual applications should consider adopting unified multi-branch style architectures rather than maintaining separate translation pipelines. For further evaluation, teams can benchmark image synthesis tools using the publicly released Animal Faces-HQ dataset and open-source models to evaluate performance under large visual domain shifts. Additional exploration into fine-tuning strategies on novel datasets is recommended prior to production deployment.
Confidence in these performance gains is high across standard aligned image domains, supported by both quantitative metrics and direct human reviews. However, performance remains validated primarily on aligned, face-centric images resized to 256 by 256 resolution. Readers should exercise caution when extrapolating these results to unaligned scenes, complex background variations, or higher-resolution visual translation tasks without further domain-specific validation.
- Paper: StarGAN: Unified Generative Adversarial Networks for Multi-domain Image-to-Image Translation, Yunjey Choi et al. (2018). StarGAN establishes the core multi-domain image-to-image translation framework with a single generator and discriminator that StarGAN v2 directly builds upon and enhances for diversity.
- Paper: Multimodal Unsupervised Image-to-Image Translation, Xun Huang et al. (2018). This paper introduces multimodal unsupervised image translation using disentangled content and style codes, a concept foundational to StarGAN v2's diverse multi-domain synthesis.
- Paper: A Style-Based Generator Architecture for Generative Adversarial Networks, Tero Karras et al. (2019). It introduces style-based modulation and latent mapping networks that inspire StarGAN v2's style code injection and mapping architecture.
- Paper: Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks, Jun-Yan Zhu et al. (2017). It formulates cycle-consistent adversarial translation for unpaired datasets, providing the underlying cyclic training principles extended to multi-domain diverse settings in StarGAN v2.
- Paper: Learning to Discover Cross-Domain Relations with Generative Adversarial Networks, Taeksoo Kim et al. (2017). This work formulates bijective cross-domain mapping and addresses mode collapse in unpaired translation, setting early foundational constraints for image-to-image translation models.
- Paper: Unsupervised Image-to-Image Translation Networks, Ming-Yu Liu et al. (2017). It introduces shared latent spaces with coupled GANs and variational autoencoders for unpaired translation, providing essential conceptual background for multi-domain mapping.
- Paper: Image-to-Image Translation with Conditional Adversarial Networks, Phillip Isola et al. (2017). Pix2Pix establishes the general conditional adversarial network framework for image-to-image translation upon which multi-domain models are constructed.
- Paper: cGANs with Projection Discriminator, Takeru Miyato et al. (2018). It provides the projection discriminator mechanism for conditioning adversarial training on domain labels, which improves multi-class and multi-domain discriminator designs.
- Paper: Analyzing and Improving the Image Quality of StyleGAN, Tero Karras et al. (2020). StyleGAN2 refines the style-based generation mechanisms utilized in modern translation architectures by eliminating droplet artifacts and stabilizing high-resolution image synthesis.
- Paper: Alias-Free Generative Adversarial Networks, Tero Karras et al. (2021). Alias-Free GAN (StyleGAN3) redesigns generative architectures to achieve continuous equivariance, directly evaluating on datasets popularized by StarGAN v2 like AFHQ.
- Paper: Efficient Geometry-aware 3D Generative Adversarial Networks, Eric R. Chan et al. (2022). EG3D extends 2D multi-domain and multi-class generative synthesis methods into high-resolution, geometry-aware 3D generation, benchmarking on StarGAN v2's AFHQ dataset.
- Paper: Palette: Image-to-Image Diffusion Models, Chitwan Saharia et al. (2021). Palette extends image-to-image translation paradigms from GAN frameworks like StarGAN v2 to multi-task conditional diffusion models.
