Toward Multimodal Image-to-Image Translation

Jun-Yan ZhuRichard ZhangDeepak PathakTrevor DarrellAlexei A. EfrosOliver WangEli Shechtman

article2017NeurIPS1,472 citations

Proposes a conditional generative framework that enforces bijective consistency between latent codes and generated images, effectively overcoming mode collapse to produce diverse, realistic outputs for ambiguous image-to-image translation tasks.

Listen

Many real-world computer vision tasks involve translating an image from one domain to another, such as converting a daytime scene to nighttime or generating a photo from an architectural sketch. These tasks are inherently ambiguous because a single input image can legitimately correspond to many different plausible outputs. Existing conditional image generation techniques tend to produce only a single deterministic result or suffer from mode collapse, where the system ignores input randomness and repeatedly generates the same few patterns. This limitation severely restricts their utility for creative and practical applications that require diverse, realistic options.

The article evaluates and demonstrates a framework, termed BicycleGAN, designed to model a full distribution of potential outputs from a single input image. The primary objective is to produce results that are simultaneously perceptually realistic, faithful to the input context, and visually diverse.

To achieve this, the article establishes a two-way, invertible mapping between the generated output images and a low-dimensional latent code. The framework was evaluated across multiple benchmark translation tasks, including edge drawings to photos, semantic labels to building facades, and aerial maps to satellite imagery. Performance was assessed through systematic testing, using human perceptual evaluations to measure photorealism and a standard deep feature distance metric to quantify output diversity across thousands of generated image pairs.

The findings show that combining latent-to-image and image-to-latent cycle constraints significantly outperforms existing approaches. In human perceptual testing on map-to-satellite translation, the BicycleGAN model achieved a 34.33% fooling rate, notably higher than baseline methods such as pix2pix with noise (27.93%) and standard conditional autoencoders (13.64%). Concurrently, the model maintained high diversity without suffering from the mode collapse seen in alternative approaches, where roughly 15% of outputs collapsed to identical results. Comparative architecture tests also confirmed that residual network encoders provide superior image reconstruction performance across tasks.

These results demonstrate that enforcing bidirectional consistency allows generative models to navigate ambiguous image tasks reliably without sacrificing visual fidelity. For organizations deploying generative image systems, this approach reduces the risk of repetitive, unrepresentative outputs and improves automated synthesis quality across commercial design and simulation pipelines.

Practitioners implementing these models should adopt residual network encoders and tune the dimensionality of the latent code based on the ambiguity of the specific dataset. The primary limitation of the study is that the latent variables do not yet map to explicitly controlled, human-interpretable attributes. Future initiatives should focus on structuring the latent space to allow users to directly manipulate specific visual parameters, such as lighting, season, or texture.

  • Paper: Image-to-Image Translation with Conditional Adversarial Networks, Phillip Isola et al. (2017). This foundational pix2pix paper establishes the paired conditional GAN framework for image-to-image translation that BicycleGAN directly builds upon to address mode collapse and multimodal distribution modeling.
  • Paper: Learning Structured Output Representation using Deep Conditional Generative Models, Kihyuk Sohn et al. (2015). This work introduces conditional variational autoencoders (CVAEs) for structured one-to-many prediction, providing the core latent variable modeling strategy combined with GAN objectives in multimodal translation.
  • Paper: Adversarial Feature Learning, Jeff Donahue et al. (2016). This paper presents Bidirectional GANs (BiGAN) for jointly learning generative mapping and latent-space encoding, inspiring the bijective consistency and encoder-generator coupling mechanisms used in the source.
  • Paper: Autoencoding beyond pixels using a learned similarity metric, Anders Boesen Lindbo Larsen et al. (2015). This work establishes the hybrid VAE-GAN framework that combines variational latent representations with adversarial training losses to generate sharp, realistic outputs.
  • Paper: Tutorial on Variational Autoencoders, Carl Doersch (2016). This tutorial details the variational inference formulations, reparameterization mechanics, and conditional VAE architectures foundational to understanding latent-space sampling in image synthesis.
  • Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). This foundational text introduces generative adversarial networks and the minimax optimization principles underlying conditional translation and discriminator objectives.
Cover for Toward Multimodal Image-to-Image Translation

Abstract

Many image-to-image translation problems are ambiguous, as a single input image may correspond to multiple possible outputs. In this work, we aim to model a \emph{distribution} of possible outputs in a conditional generative modeling setting. The ambiguity of the mapping is distilled in a low-dimensional latent vector, which can be randomly sampled at test time. A generator learns to map the given input, combined with this latent code, to the output. We explicitly encourage the connection between output and the latent code to be invertible. This helps prevent a many-to-one mapping from the latent code to the output during training, also known as the problem of mode collapse, and produces more diverse results. We explore several variants of this approach by employing different training objectives, network architectures, and methods of injecting the latent code. Our proposed method encourages bijective consistency between the latent encoding and output modes. We present a systematic comparison of our method and other variants on both perceptual realism and diversity.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Multimodal Image-to-Image Translation
  • 3.1 Baseline: pix2pix+noise (𝐳→𝐁^\mathbf{z}\rightarrow\widehat{\mathbf{B}})
  • 3.2 Conditional Variational Autoencoder GAN: cVAE-GAN (𝐁→𝐳→𝐁^\mathbf{B}\rightarrow\mathbf{z}\rightarrow\widehat{\mathbf{B}})
  • 3.3 Conditional Latent Regressor GAN: cLR-GAN (𝐳→𝐁^→𝐳^\mathbf{z}\rightarrow\widehat{\mathbf{B}}\rightarrow\widehat{\mathbf{z}})
  • 3.4 Our Hybrid Model: BicycleGAN
  • 4 Implementation Details
  • 5 Experiments
  • 5.1 Qualitative Evaluation
  • 5.2 Quantitative Evaluation
  • 6 Conclusions
  • References

Knowls

  1. Knowl 1 — BicycleGAN Framework for Multimodal Image-to-Image Translation

    model/method

    In conditional image-to-image translation, a single input image A∈A⊂RH×W×3A \in \mathcal{A} \subset \mathbb{R}^{H \times W \times 3} may correspond to multiple plausible output images B∈B⊂RH×W×3B \in \mathcal{B} \subset \mathbb{R}^{H \times W \times 3}, defining a multimodal distribution p(B∣A)p(B|A). The BicycleGAN framework distills this ambiguity into a low-dimensional stochastic latent code z∈RZ∼N(0,I)z \in \mathbb{R}^Z \sim \mathcal{N}(0, I) and learns a deterministic generator G:(A,z)→BG: (A, z) \to B together with an encoder E:B→zE: B \to z.

    To prevent mode collapse (where the generator ignores the latent vector zz and collapses to a single deterministic output), BicycleGAN explicitly enforces bijective consistency between the latent space and the output image space by jointly optimizing two complementary cycles:

    1. Image reconstruction cycle (B→z→B^B \to z \to \hat{B}): An encoder maps a ground-truth image BB into latent distribution Q(z∣B)=E(B)Q(z|B) = E(B). Latent vector zz sampled from this distribution is combined with input AA in generator G(A,z)G(A, z) to reconstruct the target image BB, regularized by a Kullback-Leibler divergence loss to align Q(z∣B)Q(z|B) with standard Gaussian prior N(0,I)\mathcal{N}(0, I).

    2. Latent reconstruction cycle (z→B^→z^z \to \hat{B} \to \hat{z}): A latent vector is sampled randomly from the prior z∼N(0,I)z \sim \mathcal{N}(0, I) and fed with AA to generator G(A,z)G(A, z) to generate an output B^\hat{B}. An encoder then reconstructs the latent vector z^=E(B^)\hat{z} = E(\hat{B}).

    At test time, diverse and realistic outputs are synthesized for any given input image AA by sampling random vectors z∼N(0,I)z \sim \mathcal{N}(0, I) and evaluating G(A,z)G(A, z).

  2. Knowl 2 — BicycleGAN Training Objective

    equation

    The BicycleGAN model jointly trains a generator GG, an encoder EE, and a discriminator DD using a combined minimax objective over paired training images (A,B)∼p(A,B)(A, B) \sim p(A, B) and latent codes z∼p(z)=N(0,I)z \sim p(z) = \mathcal{N}(0, I):

    min⁡G,Emax⁡DLBicycleGAN(G,D,E)=LGANVAE(G,D,E)+λL1VAE(G,E)+LGAN(G,D)+λlatentL1latent(G,E)+λKLLKL(E)\min_{G, E} \max_{D} \mathcal{L}_{\text{BicycleGAN}}(G, D, E) = \mathcal{L}_{\text{GAN}}^{\text{VAE}}(G, D, E) + \lambda \mathcal{L}_1^{\text{VAE}}(G, E) + \mathcal{L}_{\text{GAN}}(G, D) + \lambda_{\text{latent}} \mathcal{L}_1^{\text{latent}}(G, E) + \lambda_{\text{KL}} \mathcal{L}_{\text{KL}}(E)

    where the individual loss components are:

    • The conditional VAE adversarial loss: LGANVAE(G,D,E)=EA,B∼p(A,B)[log⁡D(A,B)]+EA,B∼p(A,B),z∼E(B)[log⁡(1−D(A,G(A,z)))]\mathcal{L}_{\text{GAN}}^{\text{VAE}}(G, D, E) = \mathbb{E}_{A, B \sim p(A, B)} [\log D(A, B)] + \mathbb{E}_{A, B \sim p(A, B), z \sim E(B)} [\log (1 - D(A, G(A, z)))]

    • The ground-truth image reconstruction ℓ1\ell_1 loss: L1VAE(G,E)=EA,B∼p(A,B),z∼E(B)[∥B−G(A,z)∥1]\mathcal{L}_1^{\text{VAE}}(G, E) = \mathbb{E}_{A, B \sim p(A, B), z \sim E(B)} [\|B - G(A, z)\|_1]

    • The standard conditional GAN adversarial loss on randomly sampled latent codes: LGAN(G,D)=EA,B∼p(A,B)[log⁡D(A,B)]+EA∼p(A),z∼p(z)[log⁡(1−D(A,G(A,z)))]\mathcal{L}_{\text{GAN}}(G, D) = \mathbb{E}_{A, B \sim p(A, B)} [\log D(A, B)] + \mathbb{E}_{A \sim p(A), z \sim p(z)} [\log (1 - D(A, G(A, z)))]

    • The latent code reconstruction ℓ1\ell_1 loss: L1latent(G,E)=EA∼p(A),z∼p(z)[∥z−E(G(A,z))∥1]\mathcal{L}_1^{\text{latent}}(G, E) = \mathbb{E}_{A \sim p(A), z \sim p(z)} [\|z - E(G(A, z))\|_1]

    • The Kullback-Leibler divergence regularizer: LKL(E)=EB∼p(B)[DKL(E(B)∥N(0,I))]\mathcal{L}_{\text{KL}}(E) = \mathbb{E}_{B \sim p(B)} [D_{\text{KL}}(E(B) \parallel \mathcal{N}(0, I))]

    The hyperparameters λ\lambda, λlatent\lambda_{\text{latent}}, and λKL\lambda_{\text{KL}} weight the relative importance of image reconstruction, latent recovery, and latent distribution regularization, respectively.

  3. Knowl 3 — Conditional Latent Regressor GAN (cLR-GAN)

    model/method

    The conditional Latent Regressor GAN (cLR-GAN) enforces a connection between the latent space and the output image space via the forward-backward cycle z→B^→z^z \to \hat{B} \to \hat{z}.

    A latent code is randomly sampled from a prior distribution z∼p(z)=N(0,I)z \sim p(z) = \mathcal{N}(0, I), and the generator produces a synthetic output B^=G(A,z)\hat{B} = G(A, z) conditioned on input image AA. An encoder EE then produces a deterministic point estimate z^=E(B^)\hat{z} = E(\hat{B}) attempting to recover the original latent vector zz. The objective function is formulated as:

    min⁡G,Emax⁡DLcLR-GAN(G,D,E)=LGAN(G,D)+λlatentL1latent(G,E)\min_{G, E} \max_{D} \mathcal{L}_{\text{cLR-GAN}}(G, D, E) = \mathcal{L}_{\text{GAN}}(G, D) + \lambda_{\text{latent}} \mathcal{L}_1^{\text{latent}}(G, E)

    where: L1latent(G,E)=EA∼p(A),z∼p(z)[∥z−E(G(A,z))∥1]\mathcal{L}_1^{\text{latent}}(G, E) = \mathbb{E}_{A \sim p(A), z \sim p(z)} [\|z - E(G(A, z))\|_1]

    Because zz is drawn at random and does not correspond to a specific ground-truth target BB, no image reconstruction loss ∥B−B^∥1\|B - \hat{B}\|_1 is used. While cLR-GAN forces GG to use the latent vector, training it in isolation can still lead to mode collapse, where approximately 15%15\% of latent vectors generate identical outputs independent of the input.

  4. Knowl 4 — Conditional Variational Autoencoder GAN (cVAE-GAN) and Variants

    model/method

    The conditional Variational Autoencoder GAN (cVAE-GAN) connects the ground-truth output to the latent space through the cycle B→z→B^B \to z \to \hat{B}. An encoder network EE maps target image BB to Gaussian parameters μ(B)\mu(B) and Σ(B)\Sigma(B), defining distribution Q(z∣B)=N(μ(B),Σ(B))Q(z|B) = \mathcal{N}(\mu(B), \Sigma(B)). Latent code zz is sampled via the reparameterization trick z=μ(B)+Σ(B)1/2⊙ϵz = \mu(B) + \Sigma(B)^{1/2} \odot \epsilon with ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I). Generator GG reconstructs target image BB from (A,z)(A, z), penalized by an ℓ1\ell_1 reconstruction loss and a conditional adversarial discriminator loss, while Q(z∣B)Q(z|B) is regularized toward N(0,I)\mathcal{N}(0, I) using KL divergence.

    Two related model variants include:

    • cAE-GAN: A deterministic baseline that drops the KL-divergence loss (setting z=E(B)z = E(B)). Because the resulting latent space is unconstrained, randomly drawing z∼N(0,I)z \sim \mathcal{N}(0, I) at test time queries unpopulated latent regions, resulting in severe visual artifacts.
    • cVAE-GAN++: An extension of cVAE-GAN that includes an additional adversarial loss LGAN(G,D)\mathcal{L}_{\text{GAN}}(G, D) evaluated on outputs generated from prior-sampled codes z∼N(0,I)z \sim \mathcal{N}(0, I). This exposes discriminator DD to samples from the prior during training but omits the latent reconstruction loss L1latent\mathcal{L}_1^{\text{latent}}.
  5. Knowl 5 — Gradient Isolation Strategy for Latent Code Reconstruction Loss

    model/method

    When optimizing the latent reconstruction loss L1latent(G,E)=EA∼p(A),z∼p(z)[∥z−E(G(A,z))∥1]\mathcal{L}_1^{\text{latent}}(G, E) = \mathbb{E}_{A \sim p(A), z \sim p(z)} [\|z - E(G(A, z))\|_1], gradients are backpropagated solely through generator GG, while the parameters of encoder EE are held fixed.

    If generator GG and encoder EE are optimized simultaneously on L1latent\mathcal{L}_1^{\text{latent}}, the two networks collude to hide imperceptible trivial signals into the synthesized image B^\hat{B} that permit EE to recover zz without GG actually producing semantically distinct output modes. By keeping EE fixed during this step, EE must rely on the feature representations learned via the VAE reconstruction branch (B→zB \to z), forcing GG to synthesize realistic visual changes that reflect the latent code.

  6. Knowl 6 — BicycleGAN Network Architecture and Training Configuration

    experimental setup

    The BicycleGAN implementation uses the following architectural and training configurations:

    • Generator GG: A U-Net architecture containing symmetric skip connections between corresponding encoder and decoder layers to preserve spatial structure between input AA and output BB.
    • Discriminator DD: Two multi-scale PatchGAN discriminators operating on image patches of sizes 70×7070 \times 70 and 140×140140 \times 140. Discriminator training uses the Least Squares GAN (LSGAN) objective. Discriminator DD is unconditional with respect to input AA (it evaluates target/generated images alone), which empirically stabilizes training.
    • Encoder EE: A convolutional neural network with residual blocks (EResNetE_{\text{ResNet}}). In the VAE branch, it outputs mean and variance; in the latent regression branch, only the predicted mean is used.
    • Latent Injection: Latent vectors z∈R8z \in \mathbb{R}^8 are spatially tiled and injected into the generator either by concatenation with input image AA (add_to_input) or by addition/concatenation to every intermediate encoder layer of GG (add_to_all).
    • Hyperparameters: Latent dimension ∣z∣=8|z| = 8; loss weights λimage=10\lambda_{\text{image}} = 10, λlatent=0.5\lambda_{\text{latent}} = 0.5, λKL=0.01\lambda_{\text{KL}} = 0.01; Adam optimizer with learning rate 0.00020.0002 and batch size 11 on 256×256256 \times 256 images.
  7. Knowl 7 — Realism versus Diversity Evaluation Across Model Variants

    data/table

    Visual realism is quantified by Amazon Mechanical Turk (AMT) fooling rate (percentage of trials where human evaluators misidentify a 1-second presentation of a generated image as real; real images achieve 50.0%50.0\%). Output diversity is quantified by average Learned Perceptual Image Patch Similarity (LPIPS) distance in AlexNet feature space across 1900 randomly sampled generated image pairs from 100 inputs on the Google maps →\to satellites dataset:

    Method AMT Fooling Rate [%] LPIPS Distance
    Random real images 50.0% 0.262 ±\pm 0.007
    pix2pix + noise 27.93 ±\pm 2.40% 0.013 ±\pm 0.000
    cAE-GAN 13.64 ±\pm 1.80% 0.204 ±\pm 0.002
    cVAE-GAN 24.93 ±\pm 2.27% 0.096 ±\pm 0.001
    cVAE-GAN++ 29.19 ±\pm 2.43% 0.098 ±\pm 0.002
    cLR-GAN 29.23 ±\pm 2.48% 0.090 ±\pm 0.002
    BicycleGAN 34.33 ±\pm 2.69% 0.110 ±\pm 0.002
    • pix2pix + noise generates realistic outputs (27.93%27.93\%) but suffers almost complete mode collapse (LPIPS 0.0130.013).
    • cAE-GAN exhibits high output variation (0.2040.204) but causes severe visual artifacts, dropping fooling rate to 13.64%13.64\%.
    • cLR-GAN exhibits mode collapse on approximately 15%15\% of images (omitted from the LPIPS computation).
    • BicycleGAN achieves the highest realism score (34.33%34.33\%) among generative models while maintaining high diversity (0.1100.110) without mode collapse.
  8. Knowl 8 — Encoder Architecture and Latent Injection Ablation Study

    data/table

    The image reconstruction performance ∥B−G(A,E(B))∥1\|B - G(A, E(B))\|_1 is evaluated across two encoder architectures (EResNetE_{\text{ResNet}} with residual blocks vs. standard ECNNE_{\text{CNN}}) and two latent code injection strategies (add_to_all vs. add_to_input) on validation datasets:

    Dataset EResNetE_{\text{ResNet}} ECNNE_{\text{CNN}}
    add_to_all add_to_input add_to_all add_to_input
    label →\to photo 0.292 ±\pm 0.058 0.292 ±\pm 0.054 0.326 ±\pm 0.066 0.339 ±\pm 0.069
    map →\to satellite 0.268 ±\pm 0.070 0.266 ±\pm 0.068 0.287 ±\pm 0.067 0.272 ±\pm 0.069

    EResNetE_{\text{ResNet}} systematically achieves lower reconstruction error than ECNNE_{\text{CNN}} across both tasks (e.g., 0.2920.292 vs. 0.3260.326 on label →\to photo). In contrast, injecting the latent vector into all intermediate generator layers (add_to_all) versus only the input layer (add_to_input) yields statistically comparable performance, demonstrating that the U-Net architecture propagates latent information throughout the network without requiring multi-layer injection.

  9. Knowl 9 — Effect of Latent Vector Dimensionality on Output Diversity

    empirical result

    The dimensionality ∣z∣|z| of the latent vector controls the capacity of BicycleGAN to represent multimodal variations:

    • Low latent dimensionality (∣z∣=2|z| = 2) limits the amount of diversity and variation that the generator can express across output modes.
    • High latent dimensionality (∣z∣=256|z| = 256) allows the network to encode rich details during reconstruction (B→zB \to z), but fitting the encoder distribution to a 256-dimensional standard Gaussian makes random sampling at inference time difficult, producing visual artifacts.
    • An intermediate dimensionality of ∣z∣=8|z| = 8 provides an optimal balance, capturing meaningful semantic variations (e.g., textures, lighting, sky conditions) while maintaining stable Gaussian sampling at inference time.

Coverage note — No substantial contributed material was omitted from the extracted knowls.

References

  1. 1.M. Arjovsky and L. Bottou. Towards principled methods for training generative adversarial networks. In ICLR, 2017.
  2. 2.A. Bansal, Y. Sheikh, and D. Ramanan. Pixelnn: Example-based image synthesis. arXiv preprint arXiv:1708.05349, 2017.
  3. 3.Q. Chen and V. Koltun. Photographic image synthesis with cascaded refinement networks. In ICCV, 2017.
  4. 4.X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel. Infogan: interpretable representation learning by information maximizing generative adversarial nets. In NIPS, 2016.
  5. 5.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016.
  6. 6.E. L. Denton, S. Chintala, A. Szlam, and R. Fergus. Deep generative image models using a laplacian pyramid of adversarial networks. In NIPS, 2015.
  7. 7.L. Dinh, J. Sohl-Dickstein, and S. Bengio. Density estimation using real nvp. In ICLR, 2017.
  8. 8.J. Donahue, P. Krahenbuhl, and T. Darrell. Adversarial feature learning. In ICLR, 2016.
  9. 9.A. Dosovitskiy and T. Brox. Generating images with perceptual similarity metrics based on deep networks. In NIPS, 2016.
  10. 10.V. Dumoulin, I. Belghazi, B. Poole, O. Mastropietro, A. Lamb, M. Arjovsky, and A. Courville. Adversarially learned inference. In ICLR, 2016.
  11. 11.A. A. Efros and T. K. Leung. Texture synthesis by non-parametric sampling. In ICCV, 1999.
  12. 12.L. A. Gatys, A. S. Ecker, and M. Bethge. Image style transfer using convolutional neural networks. In CVPR, pages 2414–2423, 2016.
  13. 13.A. Ghosh, V. Kulharia, V. Namboodiri, P. H. Torr, and P. K. Dokania. Multi-agent diverse generative adversarial networks. arXiv preprint arXiv:1704.02906, 2017.
  14. 14.I. Goodfellow. NIPS 2016 tutorial: Generative adversarial networks. arXiv preprint arXiv:1701.00160, 2016.
  15. 15.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In NIPS, 2014.
  16. 16.S. Guadarrama, R. Dahl, D. Bieber, M. Norouzi, J. Shlens, and K. Murphy. Pixcolor: Pixel recursive colorization. In BMVC, 2017.
  17. 17.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016.
  18. 18.G. E. Hinton and R. R. Salakhutdinov. Reducing the dimensionality of data with neural networks. Science, 313(5786):504–507, 2006.
  19. 19.S. Iizuka, E. Simo-Serra, and H. Ishikawa. Let there be color!: Joint end-to-end learning of global and local image priors for automatic image colorization with simultaneous classification. SIGGRAPH, 35(4), 2016.
  20. 20.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. In CVPR, 2017.
  21. 21.J. Johnson, A. Alahi, and L. Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In ECCV, 2016.
  22. 22.D. Kingma and J. Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  23. 23.D. P. Kingma and M. Welling. Auto-encoding variational bayes. In ICLR, 2014.
  24. 24.A. Krizhevsky. One weird trick for parallelizing convolutional neural networks. 2014.
  25. 25.P.-Y. Laffont, Z. Ren, X. Tao, C. Qian, and J. Hays. Transient attributes for high-level understanding and editing of outdoor scenes. SIGGRAPH, 2014.
  26. 26.A. B. L. Larsen, S. K. Sonderby, H. Larochelle, and O. Winther. Autoencoding beyond pixels using a learned similarity metric. In ICML, 2016.
  27. 27.G. Larsson, M. Maire, and G. Shakhnarovich. Learning representations for automatic colorization. In ECCV, 2016.
  28. 28.X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. P. Smolley. Least squares generative adversarial networks. In ICCV, 2017.
  29. 29.M. Mathieu, C. Couprie, and Y. LeCun. Deep multi-scale video prediction beyond mean square error. In ICLR, 2016.
  30. 30.M. Mirza and S. Osindero. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  31. 31.A. Nguyen, J. Yosinski, Y. Bengio, A. Dosovitskiy, and J. Clune. Plug & play generative networks: Conditional iterative generation of images in latent space. In CVPR, 2017.
  32. 32.A. v. d. Oord, N. Kalchbrenner, and K. Kavukcuoglu. Pixel recurrent neural networks. PMLR, 2016.
  33. 33.A. v. d. Oord, N. Kalchbrenner, O. Vinyals, L. Espeholt, A. Graves, and K. Kavukcuoglu. Conditional image generation with pixelcnn decoders. In NIPS, 2016.
  34. 34.D. Pathak, P. Krahenbuhl, J. Donahue, T. Darrell, and A. Efros. Context encoders: Feature learning by inpainting. In CVPR, 2016.
  35. 35.A. Radford, L. Metz, and S. Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks. In ICLR, 2016.
  36. 36.S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee. Generative adversarial text-to-image synthesis. In ICML, 2016.
  37. 37.O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. In MICCAI, pages 234–241. Springer, 2015.
  38. 38.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 2015.
  39. 39.T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen. Improved techniques for training gans. arXiv preprint arXiv:1606.03498, 2016.
  40. 40.P. Sangkloy, J. Lu, C. Fang, F. Yu, and J. Hays. Scribbler: Controlling deep image synthesis with sketch and color. In CVPR, 2017.
  41. 41.P. Smolensky. Information processing in dynamical systems: Foundations of harmony theory. Technical report, DTIC Document, 1986.
  42. 42.K. Sohn, X. Yan, and H. Lee. Learning structured output representation using deep conditional generative models. In NIPS, 2015.
  43. 43.P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol. Extracting and composing robust features with denoising autoencoders. In ICML, 2008.
  44. 44.J. Walker, C. Doersch, A. Gupta, and M. Hebert. An uncertain future: Forecasting from static images using variational autoencoders. In ECCV, 2016.
  45. 45.W. Xian, P. Sangkloy, J. Lu, C. Fang, F. Yu, and J. Hays. Texturegan: Controlling deep image synthesis with texture patches. In arXiv preprint arXiv:1706.02823, 2017.
  46. 46.T. Xue, J. Wu, K. Bouman, and B. Freeman. Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks. In NIPS, 2016.
  47. 47.C. Yang, X. Lu, Z. Lin, E. Shechtman, O. Wang, and H. Li. High-resolution image inpainting using multi-scale neural patch synthesis. In CVPR, 2017.
  48. 48.A. Yu and K. Grauman. Fine-grained visual comparisons with local learning. In CVPR, 2014.
  49. 49.H. Zhang, T. Xu, H. Li, S. Zhang, X. Huang, X. Wang, and D. Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In ICCV, 2017.
  50. 50.R. Zhang, P. Isola, and A. A. Efros. Colorful image colorization. In ECCV, 2016.
  51. 51.R. Zhang, J.-Y. Zhu, P. Isola, X. Geng, A. S. Lin, T. Yu, and A. A. Efros. Real-time user-guided image colorization with learned deep priors. SIGGRAPH, 2017.
  52. 52.R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang. The unreasonable effectiveness of deep features as a perceptual metric. In arXiv preprint arXiv:1801.03924, 2018.
  53. 53.J. Zhao, M. Mathieu, and Y. LeCun. Energy-based generative adversarial network. In ICLR, 2017.
  54. 54.J.-Y. Zhu, P. Krahenbuhl, E. Shechtman, and A. A. Efros. Generative visual manipulation on the natural image manifold. In ECCV, 2016.
  55. 55.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV, 2017.

Citation

MLA
Zhu, J.-Y., et al. “Toward Multimodal Image-to-Image Translation”. arXiv, 2017, http://arxiv.org/abs/1711.11586v4.
APA
Zhu, J.-Y., Zhang, R., Pathak, D., Darrell, T., Efros, A. A., Wang, O., & Shechtman, E. (2017). Toward Multimodal Image-to-Image Translation. arXiv. http://arxiv.org/abs/1711.11586v4
Chicago
Zhu, J.-Y., R. Zhang, D. Pathak, et al. 2017. “Toward Multimodal Image-to-Image Translation”. arXiv. http://arxiv.org/abs/1711.11586v4.
Harvard
Zhu, J.-Y. et al. (2017) “Toward Multimodal Image-to-Image Translation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1711.11586v4.
Vancouver
1. Zhu J-Y, Zhang R, Pathak D, Darrell T, Efros AA, Wang O, Shechtman E (2017) Toward Multimodal Image-to-Image Translation. arXiv

BibTeX

@article{zhu2017toward,
  title = {Toward Multimodal Image-to-Image Translation},
  author = {Zhu, Jun-Yan and Zhang, Richard and Pathak, Deepak and Darrell, Trevor and Efros, Alexei A. and Wang, Oliver and Shechtman, Eli},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1711.11586v4},
  eprint = {1711.11586}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors