StarGAN v2: Diverse Image Synthesis for Multiple Domains

Yunjey ChoiYoungjung UhJaejun YooJung-Woo Ha

article2019CVPR2,147 citations

Proposes a scalable image-to-image translation framework that generates diverse images across multiple visual domains within a single model and introduces the AFHQ benchmark dataset for evaluating cross-domain synthesis.

Listen

Modern computer vision applications increasingly rely on image-to-image translation to adapt visual content across distinct categories, such as changing human facial attributes or animal species. However, existing translation frameworks face a critical trade-off: methods capable of generating diverse styles typically only work between two categories and require training separate models for every category pair, while scalable multi-domain models tend to produce fixed, uniform outputs that lack style diversity. Addressing both scalability and output variety within a single system is essential for efficient and flexible generative image synthesis.

The article demonstrates that a unified generative framework, named StarGAN v2, can simultaneously achieve high visual quality, extensive style diversity, and multi-domain scalability using a single generator. It sets out to evaluate this approach across diverse image domains and benchmark its performance against leading multi-modal translation models.

The authors developed an architecture that replaces rigid domain labels with domain-specific style codes extracted from reference images or generated from random inputs. To evaluate this approach, the team conducted rigorous quantitative benchmarks and human perceptual studies using the human face dataset CelebA-HQ and a newly introduced 15,000-image Animal Faces-HQ dataset spanning cats, dogs, and wildlife. Output quality was measured via the standard Fréchet inception distance metric, while image diversity was assessed through learned perceptual image patch similarity and blind human preference surveys on Amazon Mechanical Turk.

The evaluation yielded several key findings. First, StarGAN v2 substantially improved image fidelity, achieving Fréchet inception distance scores of 13.7 on CelebA-HQ and 16.2 on Animal Faces-HQ in random-generation mode, outperforming the leading baseline by more than twofold. Second, in reference-guided translation, the model reduced error distances by approximately 1.5 times on human faces and 3.5 times on animal faces compared to the next best method. Third, the framework preserved source characteristics, such as pose and identity, while transferring fine-grained target styles like hairstyles, breeds, and eye colors. Finally, human evaluators decisively preferred StarGAN v2 over competing methods, awarding it between 68.9% and 92.1% of votes across visual quality and style-transfer categories.

These findings indicate that generative image systems do not need to trade architectural simplicity for stylistic variety. By relying on a single network for multiple categories, engineering teams can significantly lower deployment overhead, operational complexity, and the computational costs associated with training numerous pairwise models. Furthermore, the architecture successfully generalizes to unseen source distributions, making it a robust foundation for practical creative tools and automated visual content generation.

Organizations developing generative visual applications should consider adopting unified multi-branch style architectures rather than maintaining separate translation pipelines. For further evaluation, teams can benchmark image synthesis tools using the publicly released Animal Faces-HQ dataset and open-source models to evaluate performance under large visual domain shifts. Additional exploration into fine-tuning strategies on novel datasets is recommended prior to production deployment.

Confidence in these performance gains is high across standard aligned image domains, supported by both quantitative metrics and direct human reviews. However, performance remains validated primarily on aligned, face-centric images resized to 256 by 256 resolution. Readers should exercise caution when extrapolating these results to unaligned scenes, complex background variations, or higher-resolution visual translation tasks without further domain-specific validation.

  • Paper: Analyzing and Improving the Image Quality of StyleGAN, Tero Karras et al. (2020). StyleGAN2 refines the style-based generation mechanisms utilized in modern translation architectures by eliminating droplet artifacts and stabilizing high-resolution image synthesis.
  • Paper: Alias-Free Generative Adversarial Networks, Tero Karras et al. (2021). Alias-Free GAN (StyleGAN3) redesigns generative architectures to achieve continuous equivariance, directly evaluating on datasets popularized by StarGAN v2 like AFHQ.
  • Paper: Efficient Geometry-aware 3D Generative Adversarial Networks, Eric R. Chan et al. (2022). EG3D extends 2D multi-domain and multi-class generative synthesis methods into high-resolution, geometry-aware 3D generation, benchmarking on StarGAN v2's AFHQ dataset.
  • Paper: Palette: Image-to-Image Diffusion Models, Chitwan Saharia et al. (2021). Palette extends image-to-image translation paradigms from GAN frameworks like StarGAN v2 to multi-task conditional diffusion models.
Cover for StarGAN v2: Diverse Image Synthesis for Multiple Domains

Abstract

A good image-to-image translation model should learn a mapping between different visual domains while satisfying the following properties: 1) diversity of generated images and 2) scalability over multiple domains. Existing methods address either of the issues, having limited diversity or multiple models for all domains. We propose StarGAN v2, a single framework that tackles both and shows significantly improved results over the baselines. Experiments on CelebA-HQ and a new animal faces dataset (AFHQ) validate our superiority in terms of visual quality, diversity, and scalability. To better assess image-to-image translation models, we release AFHQ, high-quality animal faces with large inter- and intra-domain differences. The code, pretrained models, and dataset can be found at this https URL.

Table of Contents

  • 1 Introduction
  • 2 StarGAN v2
  • 2.1 Proposed framework
  • 2.2 Training objectives
  • 3 Experiments
  • 3.1 Analysis of individual components
  • 3.2 Comparison on diverse image synthesis
  • 4 Discussion
  • 5 Related work
  • 6 Conclusion
  • A The AFHQ dataset
  • B Training details
  • C Evaluation protocol
  • D Additional results
  • E Network architecture
  • References

Knowls

  1. Knowl 1 — StarGAN v2 Multi-Domain Architecture

    model/method

    StarGAN v2 is a scalable framework for diverse image-to-image translation across multiple visual domains using a single generator. Given a set of image domains Y\mathcal{Y} and image space X\mathcal{X}, the framework comprises four interacting modules:

    1. Generator (GG): Maps an input image x∈Xx \in \mathcal{X} and a domain-specific style code s∈Rdss \in \mathbb{R}^{d_s} (where ds=64d_s = 64) into an output image G(x,s)G(x, s). The style vector ss is injected via adaptive instance normalization (AdaIN) in the generator's upsampling layers. Because ss already encodes domain-specific style information, the generator does not require an explicit domain label input, allowing a single generator to synthesize images for all domains.

    2. Mapping Network (FF): Transforms a random latent code z∈Z⊂Rdzz \in \mathcal{Z} \subset \mathbb{R}^{d_z} (where z∼N(0,I)z \sim \mathcal{N}(0, I) and dz=16d_z = 16) and a target domain y~∈Y\tilde{y} \in \mathcal{Y} into a domain-specific style code s~=Fy~(z)\tilde{s} = F_{\tilde{y}}(z). The network is a multi-layer perceptron (MLP) with shared initial layers and domain-specific output branches (FyF_y for each domain y∈Yy \in \mathcal{Y}).

    3. Style Encoder (EE): Extracts the style code s=Ey(x)s = E_y(x) from a given reference image x∈Xx \in \mathcal{X} of domain y∈Yy \in \mathcal{Y}. It consists of a shared convolutional trunk followed by domain-specific linear heads (EyE_y for each y∈Yy \in \mathcal{Y}), enabling reference-guided translation.

    4. Discriminator (DD): A multi-task discriminator with a shared convolutional backbone and multiple linear output branches DyD_y for each y∈Yy \in \mathcal{Y}. Each branch DyD_y performs binary classification evaluating whether an input image xx is a real image of domain yy or a translated fake image G(x,s)G(x, s).

  2. Knowl 2 — StarGAN v2 Training Objectives and Loss Functions

    equation

    Let x∈Xx \in \mathcal{X} be an image with true domain y∈Yy \in \mathcal{Y}, let y~∈Y\tilde{y} \in \mathcal{Y} be a randomly selected target domain, and let z,z1,z2∈Z∼N(0,I)z, z_1, z_2 \in \mathcal{Z} \sim \mathcal{N}(0, I) be latent codes. Let s~=Fy~(z)\tilde{s} = F_{\tilde{y}}(z) be the target domain style code generated by mapping network FF. StarGAN v2 optimizes four loss objectives:

    1. Adversarial Loss: Ladv=Ex,y[log⁡Dy(x)]+Ex,y~,z[log⁡(1−Dy~(G(x,s~)))]\mathcal{L}_{\text{adv}} = \mathbb{E}_{x, y}\left[\log D_y(x)\right] + \mathbb{E}_{x, \tilde{y}, z}\left[\log \left(1 - D_{\tilde{y}}(G(x, \tilde{s}))\right)\right] where Dy(⋅)D_y(\cdot) is the discriminator output branch corresponding to domain yy.

    2. Style Reconstruction Loss: Lsty=Ex,y~,z[∥s~−Ey~(G(x,s~))∥1]\mathcal{L}_{\text{sty}} = \mathbb{E}_{x, \tilde{y}, z}\left[\|\tilde{s} - E_{\tilde{y}}(G(x, \tilde{s}))\|_1\right] which enforces that the generated image G(x,s~)G(x, \tilde{s}) reflects the target style code s~\tilde{s} as extracted by the style encoder branch Ey~E_{\tilde{y}}.

    3. Diversity Sensitive Loss: Lds=Ex,y~,z1,z2[∥G(x,s~1)−G(x,s~2)∥1]\mathcal{L}_{\text{ds}} = \mathbb{E}_{x, \tilde{y}, z_1, z_2}\left[\|G(x, \tilde{s}_1) - G(x, \tilde{s}_2)\|_1\right] where s~1=Fy~(z1)\tilde{s}_1 = F_{\tilde{y}}(z_1) and ildes2=Fy~(z2) ilde{s}_2 = F_{\tilde{y}}(z_2). Maximizing this term forces the generator to explore diverse image features. The denominator ∥z1−z2∥1\|z_1 - z_2\|_1 found in standard mode-seeking formulations is omitted to prevent gradient explosions and ensure training stability.

    4. Cycle Consistency Loss: Lcyc=Ex,y,y~,z[∥x−G(G(x,s~),s^)∥1]\mathcal{L}_{\text{cyc}} = \mathbb{E}_{x, y, \tilde{y}, z}\left[\|x - G(G(x, \tilde{s}), \hat{s})\|_1\right] where s^=Ey(x)\hat{s} = E_y(x) is the extracted style code of the original input image xx. This ensures preservation of domain-invariant characteristics (such as pose).

    Full Objective: min⁡G,F,Emax⁡DLadv+λstyLsty−λdsLds+λcycLcyc\min_{G, F, E} \max_{D} \mathcal{L}_{\text{adv}} + \lambda_{\text{sty}} \mathcal{L}_{\text{sty}} - \lambda_{\text{ds}} \mathcal{L}_{\text{ds}} + \lambda_{\text{cyc}} \mathcal{L}_{\text{cyc}} where λsty,λds,λcyc>0\lambda_{\text{sty}}, \lambda_{\text{ds}}, \lambda_{\text{cyc}} > 0 are hyperparameter balancing weights. The identical objective is also used for reference-guided training by replacing s~\tilde{s} with Ey~(xref)E_{\tilde{y}}(x_{\text{ref}}) from a target domain reference image xrefx_{\text{ref}}.

  3. Knowl 3 — Animal Faces-HQ (AFHQ) Dataset

    experimental setup

    The Animal Faces-HQ (AFHQ) dataset is an image-to-image translation dataset consisting of 15,000 high-quality animal face images at 512×512512 \times 512 resolution.

    • Domains: Contains 3 distinct domains: Cat, Dog, and Wildlife (5,000 images per domain).
    • Diversity: Each domain includes at least 8 distinct breeds or species, providing large inter-domain and intra-domain differences.
    • Data Splits: For each domain, 500 images are reserved as a held-out test set (total 1,500 test images), and the remaining 4,500 images per domain form the training set (total 13,500 training images).
    • Preprocessing: All images were collected from Flickr and Pixabay under permissive licenses, manually filtered to remove low-quality samples, and aligned horizontally and vertically such that the animal eyes are centered.
  4. Knowl 4 — StarGAN v2 Detailed Module Architectures

    model/method

    StarGAN v2 uses specific neural network layer configurations for its four modules when processing 256×256×3256 \times 256 \times 3 images:

    • Generator: Uses pre-activation residual blocks (ResBlk). It starts with a 1×11 \times 1 convolution (256×256×64256 \times 256 \times 64), followed by 4 downsampling ResBlks with average pooling and Instance Normalization (IN) down to resolution 16×16×51216 \times 16 \times 512. Next, it processes through 4 intermediate ResBlks at 16×16×51216 \times 16 \times 512 (2 with IN, 2 with AdaIN). The decoder contains 4 upsampling ResBlks with bilinear upsampling and AdaIN (channels: 512→256→128→64512 \to 256 \to 128 \to 64), ending with a 1×11 \times 1 convolution to produce 256×256×3256 \times 256 \times 3. Domain-specific style codes (s∈R64s \in \mathbb{R}^{64}) provide affine scale and shift parameters for all AdaIN layers.

    • Mapping Network: Takes a latent code z∈R16z \in \mathbb{R}^{16}. It consists of 4 shared fully-connected layers (512 units each with ReLU), followed by KK unshared domain branches. Each branch contains 3 fully-connected layers (512 units each with ReLU) and 1 linear output layer of dimension 64 (s∈R64s \in \mathbb{R}^{64}).

    • Style Encoder: A CNN consisting of a 1×11 \times 1 convolution (64 channels), 6 shared ResBlks with average pooling (feature channels: 128,256,512,512,512,512128, 256, 512, 512, 512, 512; spatial resolution reducing from 128×128128 \times 128 to 4×44 \times 4), followed by LeakyReLU, a 4×44 \times 4 convolution (1×1×5121 \times 1 \times 512), LeakyReLU, and flattening to 512 dimensions. It then forks into KK domain-specific linear layers, each mapping 512→64512 \to 64.

    • Discriminator: Follows the exact same shared convolutional trunk as the Style Encoder (6 ResBlks with average pooling down to 4×4×5124 \times 4 \times 512, LeakyReLU, 4×44 \times 4 convolution, LeakyReLU, and flattening to 512 dimensions), followed by KK unshared domain-specific linear classification heads mapping 512→1512 \to 1.

  5. Knowl 5 — StarGAN v2 Training Protocol and Hyperparameters

    experimental setup

    StarGAN v2 is trained using the following implementation parameters:

    • Optimization: Adam optimizer with β1=0\beta_1 = 0, β2=0.99\beta_2 = 0.99. Learning rates are set to 10−410^{-4} for the Generator (GG), Discriminator (DD), and Style Encoder (EE), and 10−610^{-6} for the Mapping Network (FF).
    • Adversarial Training: Non-saturating GAN loss with R1R_1 gradient penalty regularization parameter γ=1.0\gamma = 1.0.
    • Batch Size and Iterations: Batch size of 8 images trained for 100,000 iterations (~3 days on a single NVIDIA Tesla V100 GPU in PyTorch).
    • Loss Hyperparameters:
      • CelebA-HQ: λsty=1.0\lambda_{\text{sty}} = 1.0, λds=1.0\lambda_{\text{ds}} = 1.0, λcyc=1.0\lambda_{\text{cyc}} = 1.0.
      • AFHQ: λsty=1.0\lambda_{\text{sty}} = 1.0, λds=2.0\lambda_{\text{ds}} = 2.0, λcyc=1.0\lambda_{\text{cyc}} = 1.0.
      • To stabilize training, the diversity regularization weight λds\lambda_{\text{ds}} is linearly decayed to 0 over the 100,000 training iterations.
    • Weight Initialization and Evaluation: Initialized with He initialization with zero biases (except AdaIN scale biases initialized to 1.0). For evaluation, exponential moving averages (EMA) of network parameters are maintained for GG, FF, and EE.
  6. Knowl 6 — Evaluation Protocol for Multi-Domain Image Synthesis

    experimental setup

    Quantitative evaluation of visual quality and diversity is conducted using two standardized metrics on unseen test datasets:

    1. Fréchet Inception Distance (FID): Measures the distance between feature representations of real images and generated images. Features are extracted from the final average pooling layer of an ImageNet-pretrained Inception-V3 network. For every test image in a source domain, 10 translated images are generated for a target domain (using 10 randomly sampled Gaussian latent vectors for latent-guided synthesis, or 10 randomly sampled target-domain test images for reference-guided synthesis). FID is computed between all generated images and all training images of the target domain. The metric is computed across every ordered pair of domains and reported as the average.

    2. Learned Perceptual Image Patch Similarity (LPIPS): Measures the diversity among generated images using the L1L_1 distance between feature representations extracted from an ImageNet-pretrained AlexNet. For each source test image, 10 target domain images are synthesized. The average pairwise distance across all (102)=45\binom{10}{2} = 45 unique generated pairs is computed. The reported LPIPS is the mean over all source test images.

  7. Knowl 7 — Ablation Analysis of StarGAN v2 Components on CelebA-HQ

    data/table

    An ablation study evaluated the cumulative effect of each architectural and regularizing component added to baseline StarGAN on the CelebA-HQ dataset (256×256256 \times 256 resolution, male and female domains):

    Method FID ↓\downarrow LPIPS ↑\uparrow
    A Baseline StarGAN 98.4 -
    B + Multi-task discriminator 91.4 -
    C + Tuning (e.g. R1R_1 regularization, AdaIN) 80.5 -
    D + Latent code injection 32.3 0.312
    E + Replace (D) with style code 17.1 0.405
    F + Diversity regularization (StarGAN v2) 13.7 0.452
    • Baseline StarGAN (A) yields an FID of 98.4 and only produces local modifications (e.g., makeup).
    • Replacing the ACGAN discriminator with a multi-task discriminator (B) allows global structure transformation, improving FID to 91.4.
    • Incorporating R1R_1 regularization and AdaIN (C) improves stability and reduces FID to 80.5.
    • Directly injecting a Gaussian latent code with latent reconstruction loss (D) introduces multi-modality (LPIPS 0.312, FID 32.3), but latent codes fail to isolate domain-specific styles and primarily alter shared features like color.
    • Replacing direct latent injection with domain-specific style codes via the mapping network and style encoder (E) substantially improves quality (FID 17.1) and diversity (LPIPS 0.405).
    • Adding the diversity-sensitive regularization loss (F, StarGAN v2) achieves the best performance with an FID of 13.7 and LPIPS of 0.452.
  8. Knowl 8 — Latent-Guided Synthesis Performance on CelebA-HQ and AFHQ

    data/table

    Performance of StarGAN v2 compared against leading multi-modal two-domain baselines (MUNIT, DRIT, MSGAN) on latent-guided image synthesis across CelebA-HQ and AFHQ datasets (256×256256 \times 256 resolution):

    CelebA-HQ AFHQ
    Method FID ↓\downarrow LPIPS ↑\uparrow FID ↓\downarrow LPIPS ↑\uparrow
    MUNIT 31.4 0.363 41.5 0.511
    DRIT 52.1 0.178 95.6 0.326
    MSGAN 33.1 0.389 61.4 0.517
    StarGAN v2 13.7 0.452 16.2 0.450
    Real images 14.8 - 12.9 -
    • Baselines were trained pairwise for each domain combination (K(K−1)K(K-1) individual generators), whereas StarGAN v2 uses a single unified generator across all domains.
    • StarGAN v2 achieves superior visual quality with FIDs of 13.7 on CelebA-HQ and 16.2 on AFHQ, outperforming the best competing baseline (MUNIT: 31.4 on CelebA-HQ, 41.5 on AFHQ; MSGAN: 33.1 on CelebA-HQ, 61.4 on AFHQ) by more than a factor of two.
    • On CelebA-HQ, StarGAN v2 produces the highest diversity (LPIPS 0.452). On AFHQ, although baselines exhibit nominally higher LPIPS due to severe visual artifacts, StarGAN v2 achieves 0.450 diversity while maintaining photorealistic image fidelity.
  9. Knowl 9 — Reference-Guided Synthesis Performance on CelebA-HQ and AFHQ

    data/table

    Performance comparison on reference-guided image synthesis across CelebA-HQ and AFHQ datasets, where style codes are extracted from 10 randomly sampled target-domain test reference images:

    CelebA-HQ AFHQ
    Method FID ↓\downarrow LPIPS ↑\uparrow FID ↓\downarrow LPIPS ↑\uparrow
    MUNIT 107.1 0.176 223.9 0.199
    DRIT 53.3 0.311 114.8 0.156
    MSGAN 39.6 0.312 69.8 0.375
    StarGAN v2 23.8 0.388 19.8 0.432
    Real images 14.8 - 12.9 -
    • StarGAN v2 achieves an FID of 23.8 on CelebA-HQ (1.66×1.66\times lower than the best baseline MSGAN at 39.6) and an FID of 19.8 on AFHQ (3.5×3.5\times lower than MSGAN at 69.8).
    • StarGAN v2 attains the highest LPIPS diversity on both datasets (0.388 on CelebA-HQ, 0.432 on AFHQ). Competing models (MUNIT and DRIT) suffer severe mode collapse on AFHQ, leading to degraded LPIPS (0.199 and 0.156) and elevated FIDs (223.9 and 114.8).
  10. Knowl 10 — Human Perceptual Evaluation on Visual Quality and Style Reflection

    data/table

    User preference study conducted via Amazon Mechanical Turk (AMT) comparing StarGAN v2 against MUNIT, DRIT, and MSGAN across CelebA-HQ and AFHQ:

    CelebA-HQ AFHQ
    Method Quality (%) ↑\uparrow Style (%) ↑\uparrow Quality (%) ↑\uparrow Style (%) ↑\uparrow
    MUNIT 6.2 7.4 1.6 0.2
    DRIT 11.4 7.6 4.1 2.8
    MSGAN 13.5 10.1 6.2 4.9
    StarGAN v2 68.9 74.8 88.1 92.1
    • The study comprised 100 randomly generated comparison questions per task, each answered by 10 independent workers (76 total valid human annotators), with candidate methods shuffled randomly.
    • Evaluators separately judged which method produced the highest visual quality and which best stylized the input image according to the reference image.
    • StarGAN v2 received the vast majority of votes in all categories: on CelebA-HQ, 68.9% for quality and 74.8% for style reflection; on AFHQ, 88.1% for quality and 92.1% for style reflection, substantially outperforming all baselines.

Coverage note — Qualitative visual comparison figures (Figures 1, 3, 4, 5, 6, 7, 8, 9, 10) and video interpolation links were omitted as they represent qualitative visual demonstrations rather than self-contained factual or quantitative units.

References

  1. 1.A. Almahairi, S. Rajeshwar, A. Sordoni, P. Bachman, and A. Courville. Augmented cyclegan: Learning many-to-many mappings from unpaired data. In ICML, 2018. 2, 8
  2. 2.A. Anoosheh, E. Agustsson, R. Timofte, and L. Van Gool. Combogan: Unrestrained scalability for image domain translation. In CVPRW, 2018. 2
  3. 3.J. L. Ba, J. R. Kiros, and G. E. Hinton. Layer normalization. In arXiv preprint, 2016. 12
  4. 4.A. Brock, J. Donahue, and K. Simonyan. Large scale gan training for high fidelity natural image synthesis. In ICLR, 2019. 8
  5. 5.H. Chang, J. Lu, F. Yu, and A. Finkelstein. Pairedcyclegan: Asymmetric style transfer for applying and removing makeup. In CVPR, 2018. 8
  6. 6.W. Cho, S. Choi, D. K. Park, I. Shin, and J. Choo. Image-to-image translation via group-wise deep whitening-and-coloring transformation. In CVPR, 2019. 8
  7. 7.Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In CVPR, 2018. 2, 3, 4
  8. 8.J. Donahue and K. Simonyan. Large scale adversarial representation learning. In NeurIPS, 2019. 8
  9. 9.V. Dumoulin, E. Perez, N. Schucher, F. Strub, H. d. Vries, A. Courville, and Y. Bengio. Feature-wise transformations. In Distill, 2018. 4
  10. 10.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial networks. In NeurIPS, 2014. 8, 9
  11. 11.I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville. Improved training of wasserstein gans. In NeurIPS, 2017. 4
  12. 12.K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In ICCV, 2015. 9
  13. 13.K. He, X. Zhang, S. Ren, and J. Sun. Identity mappings in deep residual networks. In ECCV, 2016. 12
  14. 14.M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017. 4, 9
  15. 15.X. Huang and S. Belongie. Arbitrary style transfer in real-time with adaptive instance normalization. In ICCV, 2017. 2, 4, 12
  16. 16.X. Huang, M.-Y. Liu, S. Belongie, and J. Kautz. Multimodal unsupervised image-to-image translation. In ECCV, 2018. 2, 3, 4, 6, 7, 8, 12
  17. 17.L. Hui, X. Li, J. Chen, H. He, and J. Yang. Unsupervised multi-domain image translation with domain-specific encoders/decoders. In ICPR, 2018. 2
  18. 18.K. Hyunsu, J. Ho Young, P. Eunhyeok, and Y. Sungjoo. Tag2pix: Line art colorization using text tag with secat and changing loss. In ICCV, 2019. 8
  19. 19.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. 12
  20. 20.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial nets. In CVPR, 2017. 1, 8, 12
  21. 21.T. Karras, T. Aila, S. Laine, and J. Lehtinen. Progressive growing of GANs for improved quality, stability, and variation. In ICLR, 2018. 4, 9
  22. 22.T. Karras, S. Laine, and T. Aila. A style-based generator architecture for generative adversarial networks. In CVPR, 2019. 2, 8, 12
  23. 23.H. Kim, M. Kim, D. Seo, J. Kim, H. Park, S. Park, H. Jo, K. Kim, Y. Yang, Y. Kim, et al. Nsml: Meet the mlaas platform with a real-world case study. arXiv preprint arXiv:1810.09957, 2018. 8
  24. 24.T. Kim, M. Cha, H. Kim, J. K. Lee, and J. Kim. Learning to discover cross-domain relations with generative adversarial networks. In ICML, 2017. 3
  25. 25.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 9
  26. 26.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NeurIPS, 2012. 9
  27. 27.C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, et al. Photo-realistic single image super-resolution using a generative adversarial network. In CVPR, 2017. 8
  28. 28.H.-Y. Lee, H.-Y. Tseng, J.-B. Huang, M. K. Singh, and M.-H. Yang. Diverse image-to-image translation via disentangled representations. In ECCV, 2018. 2, 4, 6, 7, 8
  29. 29.M.-Y. Liu, T. Breuel, and J. Kautz. Unsupervised image-to-image translation networks. In NeurIPS, 2017. 8
  30. 30.M.-Y. Liu, X. Huang, A. Mallya, T. Karras, T. Aila, J. Lehtinen, and J. Kautz. Few-shot unsupervised image-to-image translation. In ICCV, 2019. 2, 4, 8
  31. 31.M. Lucic, M. Tschannen, M. Ritter, X. Zhai, O. Bachem, and S. Gelly. High-fidelity image generation with fewer labels. In ICML, 2019. 8
  32. 32.L. Ma, X. Jia, S. Georgoulis, T. Tuytelaars, and L. Van Gool. Exemplar guided unsupervised image-to-image translation with semantic consistency. In ICLR, 2019. 8
  33. 33.A. L. Maas, A. Y. Hannun, and A. Y. Ng. Rectifier nonlinearities improve neural network acoustic models. In ICML, 2013. 12
  34. 34.Q. Mao, H.-Y. Lee, H.-Y. Tseng, S. Ma, and M.-H. Yang. Mode seeking generative adversarial networks for diverse image synthesis. In CVPR, 2019. 2, 3, 4, 6, 7, 8
  35. 35.L. Mescheder, S. Nowozin, and A. Geiger. Which training methods for gans do actually converge? In ICML, 2018. 2, 4, 8, 9, 12
  36. 36.M. Mirza and S. Osindero. Conditional generative adversarial nets. In arXiv preprint, 2014. 4, 12
  37. 37.T. Miyato and M. Koyama. cGANs with projection discriminator. In ICLR, 2018. 12
  38. 38.S. Na, S. Yoo, and J. Choo. Miso: Mutual information loss with stochastic style representations for multimodal image-to-image translation. In arXiv preprint, 2019. 2
  39. 39.A. Odena, C. Olah, and J. Shlens. Conditional image synthesis with auxiliary classifier gans. In ICML, 2017. 4, 12
  40. 40.T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu. Semantic image synthesis with spatially-adaptive normalization. In CVPR, 2019. 8
  41. 41.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in pytorch. In NeurIPSW, 2017. 9
  42. 42.S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee. Generative adversarial text to image synthesis. In ICML, 2016. 12
  43. 43.N. Sung, M. Kim, H. Jo, Y. Yang, J. Kim, L. Lausen, Y. Kim, G. Lee, D. Kwak, J.-W. Ha, et al. Nsml: A machine learning platform that enables you to focus on your models. arXiv preprint arXiv:1712.05902, 2017. 8
  44. 44.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In CVPR, 2016. 9
  45. 45.D. Ulyanov, A. Vedaldi, and V. Lempitsky. Instance normalization: The missing ingredient for fast stylization. In arXiv preprint, 2016. 12
  46. 46.X. Wang, L. Bo, and L. Fuxin. Adaptive wing loss for robust face alignment via heatmap regression. In ICCV, 2019. 12
  47. 47.X. Wang, K. Yu, S. Wu, J. Gu, Y. Liu, C. Dong, Y. Qiao, and C. Change Loy. Esrgan: Enhanced super-resolution generative adversarial networks. In ECCV, 2018. 8
  48. 48.D. Yang, S. Hong, Y. Jang, T. Zhao, and H. Lee. Diversity-sensitive conditional generative adversarial networks. In ICLR, 2019. 3, 8
  49. 49.Y. Yazıcı, C.-S. Foo, S. Winkler, K.-H. Yap, G. Piliouras, and V. Chandrasekhar. The unusual effectiveness of averaging in gan training. In ICLR, 2019. 9
  50. 50.S. Yoo, H. Bahng, S. Chung, J. Lee, J. Chang, and J. Choo. Coloring with limited data: Few-shot colorization via memory augmented networks. In CVPR, 2019. 8
  51. 51.X. Yu, Y. Chen, T. Li, S. Liu, and G. Li. Multi-mapping image-to-image translation via learning disentanglement. In NeurIPS, 2019. 8
  52. 52.R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 4, 9
  53. 53.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networkss. In ICCV, 2017. 3, 8
  54. 54.J.-Y. Zhu, R. Zhang, D. Pathak, T. Darrell, A. A. Efros, O. Wang, and E. Shechtman. Toward multimodal image-to-image translation. In NeurIPS, 2017. 2, 3, 4, 8

Citation

MLA
Choi, Y., et al. “StarGAN V2: Diverse Image Synthesis for Multiple Domains”. arXiv, 2019, http://arxiv.org/abs/1912.01865v2.
APA
Choi, Y., Uh, Y., Yoo, J., & Ha, J.-W. (2019). StarGAN v2: Diverse Image Synthesis for Multiple Domains. arXiv. http://arxiv.org/abs/1912.01865v2
Chicago
Choi, Y., Y. Uh, J. Yoo, and J.-W. Ha. 2019. “StarGAN V2: Diverse Image Synthesis for Multiple Domains”. arXiv. http://arxiv.org/abs/1912.01865v2.
Harvard
Choi, Y. et al. (2019) “StarGAN v2: Diverse Image Synthesis for Multiple Domains”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1912.01865v2.
Vancouver
1. Choi Y, Uh Y, Yoo J, Ha J-W (2019) StarGAN v2: Diverse Image Synthesis for Multiple Domains. arXiv

BibTeX

@article{choi2019stargan,
  title = {StarGAN v2: Diverse Image Synthesis for Multiple Domains},
  author = {Choi, Yunjey and Uh, Youngjung and Yoo, Jaejun and Ha, Jung-Woo},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1912.01865v2},
  eprint = {1912.01865}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/