StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks

Han ZhangTao XuHongsheng LiShaoting ZhangXiaogang WangXiaolei HuangDimitris N. Metaxas

article2016ICCV2,966 citations

Proposes a two-stage generative adversarial architecture that decomposes text-to-image synthesis into sketch generation and detail refinement, utilizing conditioning augmentation to produce high-resolution, photo-realistic images from text descriptions.

Listen

The StackGAN paper addresses the longstanding difficulty of turning free-form text descriptions into high-resolution, photo-realistic images. Prior conditional GANs could produce only low-resolution outputs (typically 64×64) that lacked fine object parts and often failed entirely when scaled to 256×256 resolution because the distributions of real and generated images cease to overlap in high-dimensional pixel space.

The work set out to demonstrate that decomposing the generation task into two successive stages, combined with a simple regularization on the text-conditioning manifold, would allow stable training of 256×256 conditional GANs and produce images whose visual quality and text alignment measurably exceed those of existing single-stage models.

The authors trained and evaluated the two-stage architecture on three standard benchmarks—CUB birds, Oxford-102 flowers, and the more complex MS-COCO scenes—using both an automated inception score and blinded human rankings. They also ran controlled ablations that isolated the contribution of the stacked structure, the repeated text conditioning, and the proposed Conditioning Augmentation technique.

On CUB the model raised the inception score from 2.88 to 3.70 and improved average human rank from 2.81 to 1.37; similar relative gains appeared on Oxford-102 and COCO. Qualitatively, Stage-I produces coherent low-resolution sketches, while Stage-II consistently corrects shape and color errors and adds plausible fine details such as beaks, petals, and textures. The same text embedding can generate diverse yet semantically consistent images when noise or small perturbations are introduced, confirming that Conditioning Augmentation both stabilizes training and increases sample variety. Nearest-neighbor analysis shows the outputs are not simple memorizations of the training set.

These results indicate that text-to-image synthesis can now reach a resolution and realism level useful for downstream applications such as design visualization and automated photo editing. The staged approach also supplies a practical template for other conditional generation tasks that suffer from distribution mismatch at high resolution.

Further gains on complex, multi-object scenes will likely require richer scene-graph conditioning or additional refinement stages; the paper’s COCO results already show noticeably lower fidelity than the single-object bird and flower cases. Extending the method to higher resolutions or video would also need new techniques for temporal consistency and memory-efficient training.

The quantitative and human evaluations rest on established benchmarks and large sample sizes, giving reasonable confidence in the reported improvements for the domains tested. Performance remains sensitive to the quality of the initial Stage-I sketch, and the staged 1,200-epoch training schedule is computationally heavy; both factors should be weighed when planning deployment.

Abstract

Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing text-to-image approaches can roughly reflect the meaning of the given descriptions, but they fail to contain necessary details and vivid object parts. In this paper, we propose Stacked Generative Adversarial Networks (StackGAN) to generate 256x256 photo-realistic images conditioned on text descriptions. We decompose the hard problem into more manageable sub-problems through a sketch-refinement process. The Stage-I GAN sketches the primitive shape and colors of the object based on the given text description, yielding Stage-I low-resolution images. The Stage-II GAN takes Stage-I results and text descriptions as inputs, and generates high-resolution images with photo-realistic details. It is able to rectify defects in Stage-I results and add compelling details with the refinement process. To improve the diversity of the synthesized images and stabilize the training of the conditional-GAN, we introduce a novel Conditioning Augmentation technique that encourages smoothness in the latent conditioning manifold. Extensive experiments and comparisons with state-of-the-arts on benchmark datasets demonstrate that the proposed method achieves significant improvements on generating photo-realistic images conditioned on text descriptions.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Stacked Generative Adversarial Networks
  • 3.1 Preliminaries
  • 3.2 Conditioning Augmentation
  • 3.3 Stage-I GAN
  • 3.4 Stage-II GAN
  • 3.5 Implementation details
  • 4 Experiments
  • 4.1 Datasets and evaluation metrics
  • 4.2 Quantitative and qualitative results
  • 4.3 Component analysis
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — Stacked Generative Adversarial Networks (StackGAN) Architecture

    model/method

    Stacked Generative Adversarial Networks (StackGAN) decomposes the task of synthesizing high-resolution (256×256256 \times 256) photo-realistic images from text descriptions into two sequential stages:

    1. Stage-I GAN: Focuses on sketching the primitive shape and basic colors of the foreground object and generating background layouts from a random noise vector zz conditioned on a text embedding ϕt\phi_t. It outputs a low-resolution (64×6464 \times 64) image s0=G0(z,c^0)s_0 = G_0(z, \hat{c}_0).

    2. Stage-II GAN: Takes the low-resolution image s0s_0 and the text embedding ϕt\phi_t as inputs (omitting the noise vector zz). It rectifies defects, corrects distorted shapes or colors, and adds fine, photo-realistic details to produce a high-resolution (256×256256 \times 256) image G(s0,c^)G(s_0, \hat{c}).

    By building upon a roughly aligned low-resolution image, the model distribution generated by Stage-II has a higher probability of intersecting with the true high-resolution image distribution than directly generating 256×256256 \times 256 images from text alone.

  2. Knowl 2 — Conditioning Augmentation for Text-to-Image Synthesis

    model/method

    Conditioning Augmentation (CA) addresses the sparsity and discontinuity of high-dimensional text embedding spaces (often >100>100 dimensions) when training conditional GANs on limited text-image pairs.

    Instead of passing a fixed text embedding ϕt\phi_t deterministically into the generator, CA samples a latent conditioning variable c^\hat{c} from an independent Gaussian distribution:

    c^∼N(μ(ϕt),Σ(ϕt))\hat{c} \sim \mathcal{N}(\mu(\phi_t), \Sigma(\phi_t))

    where μ(ϕt)\mu(\phi_t) and the diagonal elements σ(ϕt)\sigma(\phi_t) of covariance matrix Σ(ϕt)\Sigma(\phi_t) are learned via fully connected layers from the text embedding ϕt\phi_t. Using the reparameterization trick, c^\hat{c} is computed as:

    c^=μ(ϕt)+σ(ϕt)⊙ϵ,ϵ∼N(0,I)\hat{c} = \mu(\phi_t) + \sigma(\phi_t) \odot \epsilon, \quad \epsilon \sim \mathcal{N}(0, I)

    To encourage smoothness over the conditioning manifold and prevent overfitting, the generator objective includes a Kullback-Leibler (KL) divergence regularization term:

    DKL(N(μ(ϕt),Σ(ϕt))∥N(0,I))D_{\text{KL}}(\mathcal{N}(\mu(\phi_t), \Sigma(\phi_t)) \parallel \mathcal{N}(0, I))

    This stochastic sampling introduces robustness to small perturbations along the text manifold and provides diversity in generated image poses and viewpoints for identical text inputs.

  3. Knowl 3 — Stage-I GAN Objectives and Network Architecture

    model/method

    In Stage-I of StackGAN, the generator G0G_0 and discriminator D0D_0 are trained by alternating the optimization of the following objectives:

    LD0=E(I0,t)∼pdata[log⁡D0(I0,ϕt)]+Ez∼pz,t∼pdata[log⁡(1−D0(G0(z,c^0),ϕt))]L_{D_0} = \mathbb{E}_{(I_0, t) \sim p_{\text{data}}} [\log D_0(I_0, \phi_t)] + \mathbb{E}_{z \sim p_z, t \sim p_{\text{data}}} [\log(1 - D_0(G_0(z, \hat{c}_0), \phi_t))]

    LG0=Ez∼pz,t∼pdata[log⁡(1−D0(G0(z,c^0),ϕt))]+λDKL(N(μ0(ϕt),Σ0(ϕt))∥N(0,I))L_{G_0} = \mathbb{E}_{z \sim p_z, t \sim p_{\text{data}}} [\log(1 - D_0(G_0(z, \hat{c}_0), \phi_t))] + \lambda D_{\text{KL}}(\mathcal{N}(\mu_0(\phi_t), \Sigma_0(\phi_t)) \parallel \mathcal{N}(0, I))

    where I0I_0 is a real low-resolution image (64×6464 \times 64), tt is the text description with pre-trained text embedding ϕt\phi_t, z∼N(0,I)z \sim \mathcal{N}(0, I) is an NzN_z-dimensional noise vector, c^0\hat{c}_0 is an NgN_g-dimensional conditioning vector sampled using Conditioning Augmentation, and λ=1\lambda = 1 is the regularization weight.

    Architecture Details:

    • Generator G0G_0: Text embedding ϕt\phi_t is mapped by a fully connected layer to mean μ0\mu_0 and diagonal covariance Σ0\Sigma_0. The sampled conditioning vector c^0\hat{c}_0 is concatenated with noise vector zz and passed through a series of upsampling blocks (nearest-neighbor upsampling followed by 3×33 \times 3 stride 1 convolution, Batch Normalization, and ReLU) to output a 64×64×364 \times 64 \times 3 image.
    • Discriminator D0D_0: Downsamples an input image to spatial dimension Md×MdM_d \times M_d using 4×44 \times 4 stride 2 convolutions. The text embedding ϕt\phi_t is compressed via a fully connected layer to dimension NdN_d and spatially replicated to form an Md×Md×NdM_d \times M_d \times N_d tensor. This text tensor is channel-wise concatenated with the downsampled image features, processed by a 1×11 \times 1 convolution, and finished with a fully connected layer outputting a scalar decision score.
  4. Knowl 4 — Stage-II GAN Objectives and Network Architecture

    model/method

    Stage-II of StackGAN conditions on the Stage-I generated image s0=G0(z,c^0)s_0 = G_0(z, \hat{c}_0) and the text embedding ϕt\phi_t to synthesize a high-resolution image (256×256256 \times 256). The noise vector zz is not included in Stage-II under the assumption that randomness is already preserved in s0s_0.

    The training objectives are:

    LD=E(I,t)∼pdata[log⁡D(I,ϕt)]+Es0∼pG0,t∼pdata[log⁡(1−D(G(s0,c^),ϕt))]L_D = \mathbb{E}_{(I, t) \sim p_{\text{data}}} [\log D(I, \phi_t)] + \mathbb{E}_{s_0 \sim p_{G_0}, t \sim p_{\text{data}}} [\log(1 - D(G(s_0, \hat{c}), \phi_t))]

    LG=Es0∼pG0,t∼pdata[log⁡(1−D(G(s0,c^),ϕt))]+λDKL(N(μ(ϕt),Σ(ϕt))∥N(0,I))L_G = \mathbb{E}_{s_0 \sim p_{G_0}, t \sim p_{\text{data}}} [\log(1 - D(G(s_0, \hat{c}), \phi_t))] + \lambda D_{\text{KL}}(\mathcal{N}(\mu(\phi_t), \Sigma(\phi_t)) \parallel \mathcal{N}(0, I))

    where II is a real high-resolution image, and c^∼N(μ(ϕt),Σ(ϕt))\hat{c} \sim \mathcal{N}(\mu(\phi_t), \Sigma(\phi_t)) is generated via Stage-II-specific fully connected layers from ϕt\phi_t.

    Architecture Details:

    • Generator GG: An encoder-decoder network. s0s_0 (64×6464 \times 64) is passed through downsampling blocks until spatial size Mg×MgM_g \times M_g. c^\hat{c} is spatially replicated to Mg×Mg×NgM_g \times M_g \times N_g and channel-wise concatenated with the downsampled image features. The concatenated features pass through residual blocks (two for 128×128128 \times 128 output, four for 256×256256 \times 256 output) to learn joint representations across image and text, followed by nearest-neighbor upsampling blocks to produce a 256×256×3256 \times 256 \times 3 image.
    • Discriminator DD: Identical structure to D0D_0 with extra downsampling blocks to accommodate the larger input image size.
  5. Knowl 5 — Matching-Aware Discriminator Training

    model/method

    To explicitly force the conditional GAN to align generated image features with semantic text descriptions, StackGAN utilizes matching-aware discriminators for both Stage-I and Stage-II.

    During training, the discriminator evaluates three categories of image-text pairs:

    1. Positive pairs: Real images paired with their corresponding (matching) text embeddings.
    2. Negative group 1: Real images paired with mismatched text embeddings.
    3. Negative group 2: Synthetic (generated) images paired with their conditioning text embeddings.

    This formulation penalizes the model not only for generating unrealistic images but also for generating realistic images that fail to reflect the semantics of the conditioning text.

  6. Knowl 6 — Quantitative Evaluation of StackGAN on Benchmark Datasets

    data/table

    StackGAN was evaluated against state-of-the-art text-to-image methods (GAN-INT-CLS and GAWWN) on CUB (birds), Oxford-102 (flowers), and MS COCO (complex scenes). Performance was measured using Inception Score (higher is better, computed on 30,000 samples) and human ranking (lower rank is better, ranked 1 to 3 by 10 human evaluators).

    Metric Dataset GAN-INT-CLS GAWWN Our StackGAN
    Inception score CUB 2.88±.042.88 \pm .04 3.62±.073.62 \pm .07 3.70±.04\mathbf{3.70 \pm .04}
    Oxford-102 2.66±.032.66 \pm .03 / 3.20±.01\mathbf{3.20 \pm .01}
    MS COCO 7.88±.077.88 \pm .07 / 8.45±.03\mathbf{8.45 \pm .03}
    Human rank CUB 2.81±.032.81 \pm .03 1.99±.041.99 \pm .04 1.37±.02\mathbf{1.37 \pm .02}
    Oxford-102 1.87±.031.87 \pm .03 / 1.13±.03\mathbf{1.13 \pm .03}
    MS COCO 1.89±.041.89 \pm .04 / 1.11±.03\mathbf{1.11 \pm .03}

    Compared to GAN-INT-CLS, StackGAN achieves a 28.47% improvement in Inception Score on CUB (from 2.88 to 3.70) and a 20.30% improvement on Oxford-102 (from 2.66 to 3.20). StackGAN also outperforms GAWWN without requiring part-location bounding boxes.

  7. Knowl 7 — Ablation Study on StackGAN Architecture Components

    data/table

    An ablation study on the CUB dataset evaluated the impact of Conditioning Augmentation (CA), stacked stages, output resolution, and re-feeding text at Stage-II on Inception Scores computed over 30,000 samples (all scaled to 299×299299 \times 299 prior to evaluation).

    Method CA Text twice Inception score
    64×6464 \times 64 Stage-I GAN no / 2.66±.032.66 \pm .03
    yes / 2.95±.022.95 \pm .02
    256×256256 \times 256 Stage-I GAN no / 2.48±.002.48 \pm .00
    yes / 3.02±.013.02 \pm .01
    128×128128 \times 128 StackGAN yes no 3.13±.033.13 \pm .03
    no yes 3.20±.033.20 \pm .03
    yes yes 3.35±.023.35 \pm .02
    256×256256 \times 256 StackGAN yes no 3.45±.023.45 \pm .02
    no yes 3.31±.033.31 \pm .03
    yes yes 3.70±.04\mathbf{3.70 \pm .04}

    Key Takeaways:

    1. Conditioning Augmentation: Removing CA from 256×256256 \times 256 StackGAN drops the score from 3.70 to 3.31; single-stage 256×256256 \times 256 GAN without CA suffers mode collapse (2.482.48).
    2. Stacked Structure: Single-stage 256×256256 \times 256 GAN with CA achieves only 3.02 compared to 3.70 for two-stage StackGAN.
    3. Re-reading Text: Omitting text input in Stage-II (relying only on Stage-I images) drops the score from 3.70 to 3.45, showing that Stage-II actively utilizes text to refine omitted details.
  8. Knowl 8 — StackGAN Training Protocol and Hyperparameters

    experimental setup

    StackGAN uses the following training configuration:

    • Dimensionality: Conditioning vector dimension Ng=128N_g = 128, noise vector dimension Nz=100N_z = 100, spatial feature map size for text fusion in Stage-II Mg=16M_g = 16, discriminator text spatial size Md=4M_d = 4, discriminator text embedding dimension Nd=128N_d = 128.
    • Resolutions: Stage-I output resolution W0=H0=64W_0 = H_0 = 64; Stage-II output resolution W=H=256W = H = 256.
    • Sequential Training: Stage-I GAN (D0D_0 and G0G_0) is trained for 600 epochs with Stage-II fixed. Then, Stage-I is fixed, and Stage-II GAN (DD and GG) is trained for another 600 epochs.
    • Optimizer: ADAM solver with batch size 64, initial learning rate of 0.0002, decayed by a factor of 2 every 100 epochs.
    • Regularization weight: λ=1\lambda = 1 for the KL divergence term.
  9. Knowl 9 — Latent Conditioning Manifold Smoothness via Sentence Embedding Interpolation

    empirical result

    Linearly interpolating between the pre-trained sentence embeddings ϕt1\phi_{t_1} and ϕt2\phi_{t_2} of two distinct text descriptions—while fixing the noise vector zz to a constant value—generates images with smooth, continuous visual transitions. The synthesized images gradually transition in color (e.g., from a completely red bird to a completely yellow bird) and object attributes (e.g., changing primary body color and wing color according to interpolated complex sentences) while maintaining plausible bird geometry. This demonstrates that StackGAN with Conditioning Augmentation learns a continuous and smooth latent data manifold rather than memorizing discrete training points.

  10. Knowl 10 — Failure Modes and Scene Complexity Limitations of StackGAN

    limitation

    StackGAN exhibits two primary limitations:

    1. Propagation of Stage-I Sketch Failures: Because Stage-II is conditioned directly on Stage-I outputs, when Stage-I GAN completely fails to synthesize plausible rough shapes or colors for the object, Stage-II GAN often cannot rectify the defect and fails to produce an acceptable high-resolution image.
    2. Complex Multi-Object Scenes: On datasets with multiple objects and complex background layouts (such as MS COCO), the synthesized image quality is visibly lower than on single-object, fine-grained datasets (such as CUB birds and Oxford-102 flowers), requiring more complex stacked architectures to model complex scenes.

Coverage note — None was omitted; all primary architectural components, losses, training regimes, benchmark comparisons, ablation experiments, interpolation analysis, and failure cases are represented.

References

  1. 1.M. Arjovsky and L. Bottou. Towards principled methods for training generative adversarial networks. In ICLR, 2017. 2
  2. 2.A. Brock, T. Lim, J. M. Ritchie, and N. Weston. Neural photo editing with introspective adversarial networks. In ICLR, 2017. 2
  3. 3.T. Che, Y. Li, A. P. Jacob, Y. Bengio, and W. Li. Mode regularized generative adversarial networks. In ICLR, 2017. 2
  4. 4.X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel. Infogan: Interpretable representation learning by information maximizing generative adversarial nets. In NIPS, 2016. 2
  5. 5.E. L. Denton, S. Chintala, A. Szlam, and R. Fergus. Deep generative image models using a laplacian pyramid of adversarial networks. In NIPS, 2015. 1, 2
  6. 6.C. Doersch. Tutorial on variational autoencoders. arXiv:1606.05908, 2016. 3
  7. 7.J. Gauthier. Conditional generative adversarial networks for convolutional face generation. Technical report, 2015. 3
  8. 8.I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio. Generative adversarial nets. In NIPS, 2014. 1, 2, 3
  9. 9.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016. 4
  10. 10.X. Huang, Y. Li, O. Poursaeed, J. Hopcroft, and S. Belongie. Stacked generative adversarial networks. In CVPR, 2017. 2, 3
  11. 11.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. 5
  12. 12.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. In CVPR, 2017. 2
  13. 13.D. P. Kingma and M. Welling. Auto-encoding variational bayes. In ICLR, 2014. 2, 3
  14. 14.A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther. Autoencoding beyond pixels using a learned similarity metric. In ICML, 2016. 3
  15. 15.C. Ledig, L. Theis, F. Huszar, J. Caballero, A. Aitken, A. Tejani, J. Totz, Z. Wang, and W. Shi. Photo-realistic single image super-resolution using a generative adversarial network. In CVPR, 2017. 2
  16. 16.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollr, and C. L. Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 5
  17. 17.E. Mansimov, E. Parisotto, L. J. Ba, and R. Salakhutdinov. Generating images from captions with attention. In ICLR, 2016. 2
  18. 18.L. Metz, B. Poole, D. Pfau, and J. Sohl-Dickstein. Unrolled generative adversarial networks. In ICLR, 2017. 2
  19. 19.M. Mirza and S. Osindero. Conditional generative adversarial nets. arXiv:1411.1784, 2014. 3
  20. 20.A. Nguyen, J. Yosinski, Y. Bengio, A. Dosovitskiy, and J. Clune. Plug & play generative networks: Conditional iterative generation of images in latent space. In CVPR, 2017. 2
  21. 21.M.-E. Nilsback and A. Zisserman. Automated flower classification over a large number of classes. In ICCVGIP, 2008. 5
  22. 22.A. Odena, C. Olah, and J. Shlens. Conditional image synthesis with auxiliary classifier gans. In ICML, 2017. 2
  23. 23.A. Radford, L. Metz, and S. Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks. In ICLR, 2016. 1, 2
  24. 24.S. Reed, Z. Akata, S. Mohan, S. Tenka, B. Schiele, and H. Lee. Learning what and where to draw. In NIPS, 2016. 1, 2, 3, 5, 6, 7
  25. 25.S. Reed, Z. Akata, B. Schiele, and H. Lee. Learning deep representations of fine-grained visual descriptions. In CVPR, 2016. 3, 5
  26. 26.S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee. Generative adversarial text-to-image synthesis. In ICML, 2016. 1, 2, 3, 5, 6
  27. 27.S. Reed, A. van den Oord, N. Kalchbrenner, V. Bapst, M. Botvinick, and N. de Freitas. Generating interpretable images with controllable structure. Technical report, 2016. 2
  28. 28.D. J. Rezende, S. Mohamed, and D. Wierstra. Stochastic backpropagation and approximate inference in deep generative models. In ICML, 2014. 2
  29. 29.T. Salimans, I. J. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen. Improved techniques for training gans. In NIPS, 2016. 2, 5
  30. 30.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In CVPR, 2016. 5
  31. 31.C. K. Snderby, J. Caballero, L. Theis, W. Shi, and F. Huszar. Amortised map inference for image super-resolution. In ICLR, 2017. 2
  32. 32.Y. Taigman, A. Polyak, and L. Wolf. Unsupervised cross-domain image generation. In ICLR, 2017. 2
  33. 33.A. van den Oord, N. Kalchbrenner, and K. Kavukcuoglu. Pixel recurrent neural networks. In ICML, 2016. 2
  34. 34.A. van den Oord, N. Kalchbrenner, O. Vinyals, L. Espeholt, A. Graves, and K. Kavukcuoglu. Conditional image generation with pixelcnn decoders. In NIPS, 2016. 2
  35. 35.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001, California Institute of Technology, 2011. 5
  36. 36.X. Wang and A. Gupta. Generative image modeling using style and structure adversarial networks. In ECCV, 2016. 2
  37. 37.X. Yan, J. Yang, K. Sohn, and H. Lee. Attribute2image: Conditional image generation from visual attributes. In ECCV, 2016. 2
  38. 38.J. Zhao, M. Mathieu, and Y. LeCun. Energy-based generative adversarial network. In ICLR, 2017. 2
  39. 39.J. Zhu, P. Krahenb ¨ uhl, E. Shechtman, and A. A. Efros. Generative visual manipulation on the natural image manifold. In ECCV, 2016. 2

Citation

MLA
Zhang, H., et al. “StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks”. 2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp. 5908–16, https://doi.org/10.1109/ICCV.2017.629.
APA
Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., & Metaxas, D. (2017). StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks. 2017 IEEE International Conference on Computer Vision (ICCV), 5908–5916. https://doi.org/10.1109/ICCV.2017.629
Chicago
Zhang, H., T. Xu, H. Li, et al. 2017. “StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks”. 2017 IEEE International Conference on Computer Vision (ICCV), 5908–16. https://doi.org/10.1109/ICCV.2017.629.
Harvard
Zhang, H. et al. (2017) “StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks”, 2017 IEEE International Conference on Computer Vision (ICCV). IEEE, pp. 5908–5916. Available at: https://doi.org/10.1109/ICCV.2017.629.
Vancouver
1. Zhang H, Xu T, Li H, Zhang S, Wang X, Huang X, Metaxas D (2017) StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks. In: 2017 IEEE International Conference on Computer Vision (ICCV). IEEE, pp 5908–5916

BibTeX

@inproceedings{Zhang_2017, title={StackGAN: Text to Photo-Realistic Image Synthesis with Stacked Generative Adversarial Networks}, url={http://dx.doi.org/10.1109/ICCV.2017.629}, DOI={10.1109/iccv.2017.629}, booktitle={2017 IEEE International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Zhang, Han and Xu, Tao and Li, Hongsheng and Zhang, Shaoting and Wang, Xiaogang and Huang, Xiaolei and Metaxas, Dimitris}, year={2017}, month=Oct, pages={5908–5916} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE