Generative Adversarial Text to Image Synthesis

Scott ReedZeynep AkataXinchen YanLajanugen LogeswaranBernt SchieleHonglak Lee

article2016ICML3,451 citations

Introduces a conditional generative adversarial network architecture that bridges text representations and convolutional generators to synthesize realistic images directly from natural language descriptions.

Listen

Researchers developed a conditional generative adversarial network that translates single-sentence human descriptions directly into 64-by-64 pixel images. The work addresses the long-standing difficulty of turning flexible natural-language descriptions into visual output without relying on hand-crafted attributes or class labels. Automatic image synthesis from text would support applications in design, content creation, and data augmentation, yet prior systems produced implausible or incoherent results even for narrow domains such as birds and flowers.

The authors set out to demonstrate that a single end-to-end model could learn both a discriminative text representation and a generator capable of producing visually plausible images that match the given description, including descriptions from categories never seen during training. They combined a pre-trained character-level convolutional-recurrent text encoder with a deep convolutional GAN whose generator and discriminator both receive the text embedding. Two training modifications were introduced: a matching-aware discriminator that penalizes realistic images paired with mismatched text, and a manifold-interpolation regularizer that encourages the generator to produce coherent images from synthetic text embeddings lying between training examples. Experiments used the Caltech-UCSD Birds and Oxford-102 Flowers datasets with five captions per image, plus a subset of MS-COCO for broader scenes; training and test categories were kept disjoint to evaluate zero-shot generalization.

The strongest variant, which combined both modifications, generated images that human observers found plausible and that correctly reflected color, shape, and part attributes described in the captions. On birds, interpolation regularization proved essential for visual quality; on flowers, all tested variants succeeded more readily. The model also separated content (captured by the text embedding) from style factors such as pose and background color (captured by the noise vector), enabling style transfer from a query photograph onto new text descriptions. On MS-COCO the same architecture produced sharp images that roughly matched multi-object captions, although scene coherence remained limited.

These results show that conditional GANs can move beyond class-label conditioning to open-ended text, removing the need for expensive attribute annotations while still supporting zero-shot synthesis. The approach therefore lowers the barrier to creating visual content from ordinary language and offers a practical route to controllable image generation. For deployment, organizations would still need higher-resolution output and better handling of complex scenes; the authors note that further scaling of the generator and incorporation of hierarchical structure are logical next steps. The reported findings rest on qualitative inspection and limited quantitative checks on style disentanglement; performance on entirely new visual domains or longer, more compositional text remains untested.

  • Paper: Conditional Generative Adversarial Nets, Mehdi Mirza et al. (2014). Its conditional-GAN formulation—conditioning both generator and discriminator on auxiliary information—is the core framework this paper adapts from class labels to sentence embeddings.
  • Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). The original GAN framework establishes the generator–discriminator competition that this paper uses as the basis for conditional text-to-image synthesis.
Cover for Generative Adversarial Text to Image Synthesis

Abstract

Automatic synthesis of realistic images from text would be interesting and useful, but current AI systems are still far from this goal. However, in recent years generic and powerful recurrent neural network architectures have been developed to learn discriminative text feature representations. Meanwhile, deep convolutional generative adversarial networks (GANs) have begun to generate highly compelling images of specific categories, such as faces, album covers, and room interiors. In this work, we develop a novel deep architecture and GAN formulation to effectively bridge these advances in text and image model- ing, translating visual concepts from characters to pixels. We demonstrate the capability of our model to generate plausible images of birds and flowers from detailed text descriptions.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Background
  • 3.1. Generative adversarial networks
  • 3.2. Deep symmetric structured joint embedding
  • 4. Method
  • 4.1. Network architecture
  • 4.2. Matching-aware discriminator (GAN-CLS)
  • 4.3. Learning with manifold interpolation (GAN-INT)
  • 4.4. Inverting the generator for style transfer
  • 5. Experiments
  • 5.1. Qualitative results
  • 5.2. Disentangling style and content
  • 5.3. Pose and background style transfer
  • 5.4. Sentence interpolation
  • 5.5. Beyond birds and flowers
  • 6. Conclusions
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Text-Conditional Generative Adversarial Network Architecture

    model/method

    The text-conditional Generative Adversarial Network (GAN) synthesizes images directly from natural language descriptions by conditioning both the generator GG and the discriminator DD on deep text feature embeddings.

    Let tt denote a text description, ϕ(t)∈RT\phi(t) \in \mathbb{R}^T its text embedding (with dimension T=1024T = 1024), z∈RZz \in \mathbb{R}^Z a latent noise vector sampled from a standard normal distribution N(0,IZ)\mathcal{N}(0, I_Z) with dimension Z=100Z = 100, and x∈RDx \in \mathbb{R}^D an image of dimension D=64×64×3D = 64 \times 64 \times 3.

    • Generator Network G:RZ×RT→RDG: \mathbb{R}^Z \times \mathbb{R}^T \to \mathbb{R}^D: The text embedding ϕ(t)\phi(t) is projected to a 128-dimensional vector using a fully connected layer followed by a leaky ReLU activation. This representation is concatenated along the channel dimension with the noise vector zz. The concatenated vector is then passed through a series of deconvolutional (transposed convolutional) layers with spatial batch normalization and ReLU activations to synthesize an image x^=G(z,ϕ(t))\hat{x} = G(z, \phi(t)).

    • Discriminator Network D:RD×RT→[0,1]D: \mathbb{R}^D \times \mathbb{R}^T \to [0, 1]: The discriminator processes an input image (real xx or synthetic x^\hat{x}) through multiple stride-2 convolutional layers with spatial batch normalization and leaky ReLU activations. When the spatial resolution of the feature maps reaches 4×44 \times 4, the text embedding ϕ(t)\phi(t) (separately projected to 128 dimensions via a fully connected layer and rectified) is spatially replicated over the 4×44 \times 4 grid and concatenated along the channel/depth dimension with the image feature maps. A 1×11 \times 1 convolution followed by rectification and a final 4×44 \times 4 convolution produces the real/fake score D(x,ϕ(t))∈[0,1]D(x, \phi(t)) \in [0, 1].

  2. Knowl 2 — Matching-Aware Discriminator Algorithm

    algorithm

    Standard conditional GAN training only distinguishes real image-text pairs from synthetic image-text pairs, leaving the discriminator without an explicit training signal to penalize realistic images that do not match the conditioning text. The matching-aware GAN (GAN-CLS) introduces mismatched real pairs (x,t^)(x, \hat{t}) as an explicit third input type that the discriminator must learn to score as fake.

    Given real images xx, matching text descriptions tt, and mismatched text descriptions t^\hat{t}:

    • sr=D(x,ϕ(t))s_r = D(x, \phi(t)) is the discriminator score for a real image with matching text.
    • sw=D(x,ϕ(t^))s_w = D(x, \phi(\hat{t})) is the discriminator score for a real image with mismatched text.
    • sf=D(x^,ϕ(t))s_f = D(\hat{x}, \phi(t)) is the discriminator score for a generated synthetic image x^=G(z,ϕ(t))\hat{x} = G(z, \phi(t)) with matching text.

    The training losses are given by: LD=−log⁡(sr)−12(log⁡(1−sw)+log⁡(1−sf))\mathcal{L}_D = -\log(s_r) - \frac{1}{2}\left(\log(1 - s_w) + \log(1 - s_f)\right) LG=−log⁡(sf)\mathcal{L}_G = -\log(s_f)

    The training routine using minibatch stochastic gradient descent proceeds as follows:

    Input: Minibatch of training images xx, matching text tt, mismatching text t^\hat{t}, text encoder ϕ\phi, generator GG, discriminator DD, learning rate α\alpha, number of batch steps SS.
    for n=1n = 1 to SS do
        h←ϕ(t)h \leftarrow \phi(t)
        h^←ϕ(t^)\hat{h} \leftarrow \phi(\hat{t})
        z∼N(0,IZ)z \sim \mathcal{N}(0, I_Z)
        x^←G(z,h)\hat{x} \leftarrow G(z, h)
        sr←D(x,h)s_r \leftarrow D(x, h)
        sw←D(x,h^)s_w \leftarrow D(x, \hat{h})
        sf←D(x^,h)s_f \leftarrow D(\hat{x}, h)
        LD←−log⁡(sr)−0.5⋅(log⁡(1−sw)+log⁡(1−sf))\mathcal{L}_D \leftarrow -\log(s_r) - 0.5 \cdot (\log(1 - s_w) + \log(1 - s_f))
        D←D−α∂LD∂DD \leftarrow D - \alpha \frac{\partial \mathcal{L}_D}{\partial D}
        LG←−log⁡(sf)\mathcal{L}_G \leftarrow -\log(s_f)
        G←G−α∂LG∂GG \leftarrow G - \alpha \frac{\partial \mathcal{L}_G}{\partial G}
    end for
  3. Knowl 3 — Text Manifold Interpolation Regularizer

    model/method

    To regularize the generator and expand the variety of conditioning text vectors without requiring additional human labeling, the GAN-INT method trains the generator on synthetic text embeddings formed by linearly interpolating between pairs of training embeddings.

    Given two training captions with embeddings t1,t2∼pdatat_1, t_2 \sim p_{\text{data}}, an interpolated embedding is defined as: t~=βt1+(1−β)t2\tilde{t} = \beta t_1 + (1 - \beta) t_2 where β∈[0,1]\beta \in [0, 1] (set to β=0.5\beta = 0.5 in practice). The text descriptions t1t_1 and t2t_2 can come from different images and different categories.

    Because interpolated embeddings t~\tilde{t} lack ground-truth real images, the discriminator DD cannot be trained on real pairs using t~\tilde{t}. However, since DD evaluates whether an image matches the provided text embedding, the generator GG is regularized by minimizing the adversarial objective on interpolated embeddings: LINT=Et1,t2∼pdata,z∼N(0,IZ)[log⁡(1−D(G(z,βt1+(1−β)t2)))]\mathcal{L}_{\text{INT}} = \mathbb{E}_{t_1, t_2 \sim p_{\text{data}}, z \sim \mathcal{N}(0, I_Z)} \left[ \log(1 - D(G(z, \beta t_1 + (1 - \beta) t_2))) \right]

    This term encourages the generator GG to fill in gaps on the data manifold between observed training points.

  4. Knowl 4 — Inverting the Generator for Visual Style Transfer

    model/method

    In text-conditional image synthesis, the conditioning text embedding ϕ(t)\phi(t) captures the semantic content (such as bird color patterns or flower anatomy), while the latent noise vector zz captures stylistic factors of variation (such as pose, viewpoint, and background). Visual style from a query image xx can be transferred onto a text description tt by inverting the generator GG using a style encoder network S:RD→RZS: \mathbb{R}^D \to \mathbb{R}^Z.

    The style encoder SS is trained on synthetic generator outputs to regress the input noise vector zz via mean squared error: Lstyle=Et∼pdata,z∼N(0,IZ)[∥z−S(G(z,ϕ(t)))∥22]\mathcal{L}_{\text{style}} = \mathbb{E}_{t \sim p_{\text{data}}, z \sim \mathcal{N}(0, I_Z)} \left[ \| z - S(G(z, \phi(t))) \|_2^2 \right]

    Once SS and GG are trained, style transfer from a query image xx onto a target text description tt is performed in two steps:

    1. Predict the style vector: s=S(x)s = S(x)
    2. Generate the stylized image: x^=G(s,ϕ(t))\hat{x} = G(s, \phi(t))

    This preserves the visual orientation and background style of xx while rendering the attributes specified in tt.

  5. Knowl 5 — Deep Hybrid Character-Level Convolutional-Recurrent Text Encoder

    model/method

    To extract visually discriminative text embeddings robust to typos and out-of-vocabulary words, a hybrid character-level convolutional-recurrent neural network (char-CNN-RNN) is utilized as the text encoder ϕ(t)\phi(t).

    The text encoder ϕ\phi is learned jointly with an image encoder φ(v)\varphi(v) (e.g., GoogLeNet) by optimizing a symmetric structured joint embedding loss on training triplets {(vn,tn,yn)}n=1N\{(v_n, t_n, y_n)\}_{n=1}^N of images vnv_n, text descriptions tnt_n, and class labels yn∈Yy_n \in \mathcal{Y}: 1N∑n=1N(Δ(yn,fv(vn))+Δ(yn,ft(tn)))\frac{1}{N} \sum_{n=1}^N \left( \Delta(y_n, f_v(v_n)) + \Delta(y_n, f_t(t_n)) \right) where Δ\Delta is 0-1 loss, and the image-to-text and text-to-image classifiers are defined as: fv(v)=arg⁡max⁡y∈YEt∼T(y)[φ(v)Tϕ(t)]f_v(v) = \arg\max_{y \in \mathcal{Y}} \mathbb{E}_{t \sim \mathcal{T}(y)} \left[ \varphi(v)^T \phi(t) \right] ft(t)=arg⁡max⁡y∈YEv∼V(y)[φ(v)Tϕ(t)]f_t(t) = \arg\max_{y \in \mathcal{Y}} \mathbb{E}_{v \sim \mathcal{V}(y)} \left[ \varphi(v)^T \phi(t) \right] where T(y)\mathcal{T}(y) is the set of text descriptions for class yy, and V(y)\mathcal{V}(y) is the set of images for class yy. The resulting 1,024-dimensional text representation ϕ(t)\phi(t) is pre-trained to maximize compatibility with the visual features of corresponding categories.

  6. Knowl 6 — Experimental Setup for Class-Disjoint Text-to-Image Synthesis

    experimental setup

    The text-conditional GAN framework is evaluated on fine-grained and general image synthesis tasks:

    • Datasets:

      • Caltech-UCSD Birds (CUB): Contains 11,788 bird images across 200 species, partitioned into 150 train+validation classes and 50 disjoint zero-shot test classes. Each image has 5 human-written captions.
      • Oxford-102 Flowers: Contains 8,189 flower images across 102 categories, partitioned into 82 train+validation classes and 20 disjoint zero-shot test classes. Each image has 5 human-written captions.
      • MS COCO: General domain dataset featuring complex scenes with multiple objects and diverse backgrounds.
    • Training Details:

      • Training image resolution is 64×64×364 \times 64 \times 3.
      • Text descriptions are encoded into 1,024-dimensional embeddings and projected to 128 dimensions in both the generator and discriminator.
      • Latent noise vectors are sampled from a 100-dimensional standard normal distribution z∼N(0,I100)z \sim \mathcal{N}(0, I_{100}).
      • Optimization uses Adam with a base learning rate of 0.0002 and momentum β1=0.5\beta_1 = 0.5.
      • Minibatch size is 64, trained for 600 epochs with alternating generator and discriminator updates.
  7. Knowl 7 — Style-Content Disentanglement and Style Verification on CUB

    empirical result

    Disentanglement between semantic text content and style factors (pose and background) captured by the latent vector zz was quantitatively evaluated on the CUB dataset using generator inversion:

    • Verification Tasks:

      • Pose Verification: Images were grouped into 100 clusters via K-means on 6 anatomical keypoint coordinates (beak, belly, breast, crown, forehead, and tail).
      • Background Color Verification: Images were grouped into 100 clusters via K-means on mean background RGB color.
      • For both tasks, style vectors s=S(x)s = S(x) were predicted by the inverted generator network SS, and cosine similarities between same-cluster versus different-cluster pairs were evaluated using 5-fold Area Under the ROC Curve (AU-ROC).
    • Results:

      • Text feature baselines yielded near-chance AU-ROC, demonstrating that captions do not encode background or pose details.
      • Models trained with the manifold interpolation regularizer (GAN-INT and GAN-INT-CLS) achieved substantially higher AU-ROC on both pose verification and background verification than the baseline GAN and GAN-CLS models, confirming that the latent noise vector zz captures style factors independently of text content.
  8. Knowl 8 — Zero-Shot Synthesis and Latent Space Interpolation Properties

    empirical result

    Empirical evaluation on fine-grained benchmarks demonstrated the following properties:

    • Zero-Shot Generalization: On unseen categories in the CUB test split, baseline GAN and GAN-CLS captured approximate colors but generated unrealistic bird shapes. Models incorporating manifold interpolation (GAN-INT and GAN-INT-CLS) reliably synthesized plausible 64×6464 \times 64 images matching the textual attributes. On Oxford-102 Flowers, all variants synthesized plausible images, with regularized variants producing higher class-consistent morphology.
    • Text Manifold Interpolation: Linearly interpolating between two text embeddings while keeping the latent noise vector zz fixed produced smooth transitions in semantic visual attributes (e.g., color changing continuously from blue to red) while keeping pose and background invariant.
    • Noise Space Interpolation: Linearly interpolating between two latent noise vectors z1,z2z_1, z_2 while keeping the text embedding fixed produced smooth transitions in bird pose, background, and lighting while keeping fine-grained bird attributes constant.
  9. Knowl 9 — Complex Multi-Object Scene Synthesis and Temporal Text Limitations

    limitation

    The text-conditional DC-GAN framework exhibits key limitations when extended beyond single-object domains:

    • Multi-Object Incoherence: When trained on complex multi-object scenes from the MS COCO dataset, the model produces sharp textures and matching color schemes but frequently fails to generate coherent structural geometry (e.g., human figures in sports scenes appear as indistinct blobs lacking articulated limbs or faces).
    • Absence of Hierarchical Compositionality: The architecture lacks explicit hierarchical mechanisms to decompose complex scenes into individual objects, distinct background layers, and spatial relationships.
    • Sensitivity to Fine Text Variations: Compared to iterative recurrent models with attention (such as AlignDRAW), the single-step feed-forward GAN generator is less sensitive to fine single-word modifications in conditioning text prompts, indicating that explicit temporal or spatial attention mechanisms are needed to capture fine-grained prompt variations.

Coverage note — None was omitted; the extracted knowls comprehensively cover the conditional GAN architecture, the GAN-CLS matching-aware discriminator, the GAN-INT manifold interpolation regularizer, the style inversion encoder, text encoder training, experimental setup, quantitative disentanglement results, qualitative zero-shot synthesis findings, and stated limitations.

References

  1. 1.Akata, Z., Reed, S., Walter, D., Lee, H., and Schiele, B. Evaluation of Output Embeddings for Fine-Grained Image Classification. In CVPR, 2015.
  2. 2.Ba, J. and Kingma, D. Adam: A method for stochastic optimization. In ICLR, 2015.
  3. 3.Bengio, Y., Mesnil, G., Dauphin, Y., and Rifai, S. Better mixing via deep representations. In ICML, 2013.
  4. 4.Denton, E. L., Chintala, S., Fergus, R., et al. Deep generative image models using a laplacian pyramid of adversarial networks. In NIPS, 2015.
  5. 5.Donahue, J., Hendricks, L. A., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., and Darrell, T. Long-term recurrent convolutional networks for visual recognition and description. In CVPR, 2015.
  6. 6.Dosovitskiy, A., Tobias Springenberg, J., and Brox, T. Learning to generate chairs with convolutional neural networks. In CVPR, 2015.
  7. 7.Farhadi, A., Endres, I., Hoiem, D., and Forsyth, D. Describing objects by their attributes. In CVPR, 2009.
  8. 8.Fu, Y., Hospedales, T. M., Xiang, T., Fu, Z., and Gong, S. Transductive multi-view embedding for zero-shot recognition and annotation. In ECCV, 2014.
  9. 9.Gauthier, J. Conditional generative adversarial nets for convolutional face generation. Technical report, 2015.
  10. 10.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial nets. In NIPS, 2014.
  11. 11.Gregor, K., Danihelka, I., Graves, A., Rezende, D., and Wierstra, D. Draw: A recurrent neural network for image generation. In ICML, 2015.
  12. 12.Hochreiter, S. and Schmidhuber, J. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  13. 13.Ioffe, S. and Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015.
  14. 14.Karpathy, A. and Li, F. Deep visual-semantic alignments for generating image descriptions. In CVPR, 2015.
  15. 15.Kiros, R., Salakhutdinov, R., and Zemel, R. S. Unifying visual-semantic embeddings with multimodal neural language models. In ACL, 2014.
  16. 16.Kumar, N., Berg, A. C., Belhumeur, P. N., and Nayar, S. K. Attribute and simile classifiers for face verification. In ICCV, 2009.
  17. 17.Lampert, C. H., Nickisch, H., and Harmeling, S. Attribute-based classification for zero-shot visual object categorization. TPAMI, 36(3):453–465, 2014.
  18. 18.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In ECCV. 2014.
  19. 19.Mansimov, E., Parisotto, E., Ba, J. L., and Salakhutdinov, R. Generating images from captions with attention. ICLR, 2016.
  20. 20.Mao, J., Xu, W., Yang, Y., Wang, J., and Yuille, A. Deep captioning with multimodal recurrent neural networks (m-rnn). ICLR, 2015.
  21. 21.Mirza, M. and Osindero, S. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  22. 22.Ngiam, J., Khosla, A., Kim, M., Nam, J., Lee, H., and Ng, A. Y. Multimodal deep learning. In ICML, 2011.
  23. 23.Parikh, D. and Grauman, K. Relative attributes. In ICCV, 2011.
  24. 24.Radford, A., Metz, L., and Chintala, S. Unsupervised representation learning with deep convolutional generative adversarial networks. 2016.
  25. 25.Reed, S., Sohn, K., Zhang, Y., and Lee, H. Learning to disentangle factors of variation with manifold interaction. In ICML, 2014.
  26. 26.Reed, S., Zhang, Y., Zhang, Y., and Lee, H. Deep visual analogy-making. In NIPS, 2015.
  27. 27.Reed, S., Akata, Z., Lee, H., and Schiele, B. Learning deep representations for fine-grained visual descriptions. In CVPR, 2016.
  28. 28.Ren, M., Kiros, R., and Zemel, R. Exploring models and data for image question answering. In NIPS, 2015.
  29. 29.Sohn, K., Shang, W., and Lee, H. Improved multimodal deep learning with variation of information. In NIPS, 2014.
  30. 30.Srivastava, N. and Salakhutdinov, R. R. Multimodal learning with deep boltzmann machines. In NIPS, 2012.
  31. 31.Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. Going deeper with convolutions. In CVPR, 2015.
  32. 32.Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. Show and tell: A neural image caption generator. In CVPR, 2015.
  33. 33.Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. The caltech-ucsd birds-200-2011 dataset. 2011.
  34. 34.Wang, P., Wu, Q., Shen, C., Hengel, A. v. d., and Dick, A. Explicit knowledge-based reasoning for visual question answering. arXiv preprint arXiv:1511.02570, 2015.
  35. 35.Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., and Bengio, Y. Show, attend and tell: Neural image caption generation with visual attention. In ICML, 2015.
  36. 36.Yan, X., Yang, J., Sohn, K., and Lee, H. Attribute2image: Conditional image generation from visual attributes. arXiv preprint arXiv:1512.00570, 2015.
  37. 37.Yang, J., Reed, S., Yang, M.-H., and Lee, H. Weakly-supervised disentangling with recurrent transformations for 3d view synthesis. In NIPS, 2015.
  38. 38.Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In ICCV, 2015.

Citation

MLA
Reed, S., et al. “Generative Adversarial Text to Image Synthesis”. arXiv, 2016, http://arxiv.org/abs/1605.05396v2.
APA
Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., & Lee, H. (2016). Generative Adversarial Text to Image Synthesis. arXiv. http://arxiv.org/abs/1605.05396v2
Chicago
Reed, S., Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee. 2016. “Generative Adversarial Text to Image Synthesis”. arXiv. http://arxiv.org/abs/1605.05396v2.
Harvard
Reed, S. et al. (2016) “Generative Adversarial Text to Image Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1605.05396v2.
Vancouver
1. Reed S, Akata Z, Yan X, Logeswaran L, Schiele B, Lee H (2016) Generative Adversarial Text to Image Synthesis. arXiv

BibTeX

@article{reed2016generative,
  title = {Generative Adversarial Text to Image Synthesis},
  author = {Reed, Scott and Akata, Zeynep and Yan, Xinchen and Logeswaran, Lajanugen and Schiele, Bernt and Lee, Honglak},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1605.05396v2},
  eprint = {1605.05396}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors