DualGAN: Unsupervised Dual Learning for Image-to-Image Translation

Zili YiHao ZhangPing TanMinglun Gong

article2017ICCV2,121 citations

Proposes DualGAN, an unsupervised framework that leverages closed-loop dual learning to translate images between domains without paired training data, achieving results competitive with fully supervised models.

Listen

Translating images across visual styles and modalities—such as converting sketches to photographs, changing daylight scenes to night, or altering artistic styles—historically requires deep neural networks trained on thousands of precisely matched image pairs. However, collecting labeled pairs that depict the exact same content under identical alignments is labor-intensive, costly, and frequently impossible in real-world scenarios.

The article develops and evaluates DualGAN, a framework designed to perform general-purpose image-to-image translation using two independent sets of unlabeled, unpaired images without requiring pre-trained representations or domain-specific supervision.

The authors designed a dual-learning architecture where two generator networks form a closed translation loop: one translates an image from the first style domain to the second, and the other translates it back to reconstruct the original input. The system measures the differences between the original and reconstructed images to optimize the models alongside adversarial discriminators. The authors tested this approach across diverse translation benchmarks, including photo-sketch pairs, daylight-to-night scene transitions, architectural facade segmentations, map conversions, material synthesis, and artistic painting style transfers. Human perception and visual realism were systematically quantified using large-scale participant evaluation studies.

The key findings demonstrate that DualGAN consistently outperforms conventional unsupervised generative adversarial networks across all benchmarks, generating significantly sharper outputs with fewer visual artifacts. In human realness evaluations, DualGAN surpassed a standard generative baseline by wide margins, scoring 2.42 out of 4 versus 0.13 for day-to-night conversions and 1.87 versus 1.04 for sketch-to-photo translations. Furthermore, DualGAN outperformed fully supervised models on daylight-to-night and sketch-to-photo conversions, proving more resilient against misalignments that frequently degrade paired training data. However, for semantic labeling tasks such as parsing building facades or converting aerial photographs into maps, DualGAN lagged behind supervised models, achieving a facade pixel accuracy of 0.27 compared to 0.54 for supervised networks.

These results show that organizations can successfully bypass expensive, manual data-labeling pipelines for texture-, style-, and appearance-driven image translation tasks, dramatically reducing project setup costs and development timelines. While DualGAN excels at visual styling without paired data, purely unsupervised frameworks struggle to deduce arbitrary semantic associations (such as linking specific colors to precise categorical map labels) without explicit supervision.

Decision-makers should deploy unsupervised dual-learning pipelines for applications centered on visual enhancement, stylization, and material rendering where unpaired data is abundant. For mission-critical tasks requiring exact semantic parsing or regulatory categorization, teams should avoid relying purely on unsupervised models. Future research and implementation should explore hybrid approaches, combining dual-learning models with a small set of labeled data as a initialization step to close the semantic accuracy gap.

Cover for DualGAN: Unsupervised Dual Learning for Image-to-Image Translation

Abstract

Conditional Generative Adversarial Networks (GANs) for cross-domain image-to-image translation have made much progress recently. Depending on the task complexity, thousands to millions of labeled image pairs are needed to train a conditional GAN. However, human labeling is expensive, even impractical, and large quantities of data may not always be available. Inspired by dual learning from natural language translation, we develop a novel dual-GAN mechanism, which enables image translators to be trained from two sets of unlabeled images from two domains. In our architecture, the primal GAN learns to translate images from domain U to those in domain V, while the dual GAN learns to invert the task. The closed loop made by the primal and dual tasks allows images from either domain to be translated and then reconstructed. Hence a loss function that accounts for the reconstruction error of images can be used to train the translators. Experiments on multiple image translation tasks with unlabeled data show considerable performance gain of DualGAN over a single GAN. For some tasks, DualGAN can even achieve comparable or slightly better results than conditional GAN trained on fully labeled data.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 Method
  • 3.1 Objective
  • 3.2 Network configuration
  • 3.3 Training procedure
  • 4 Experimental results and evaluation
  • 5 Qualitative evaluation
  • 5.1 Quantitative evaluation
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — DualGAN Framework for Unsupervised Image-to-Image Translation

    model/method

    DualGAN is an unsupervised deep learning framework designed to perform cross-domain image-to-image translation without requiring paired or labeled training examples. Given two sets of unpaired images sampled from domain UU and domain VV, DualGAN establishes a primal-dual closed loop consisting of two generative adversarial networks:

    1. Primal GAN: Learns a primal generator GA:U→VG_A: U \to V to map an image u∈Uu \in U to the domain VV, and an adversarial discriminator DAD_A trained to distinguish generated fake samples GA(u,z)G_A(u, z) from authentic images v∈Vv \in V.
    2. Dual GAN: Learns a dual generator GB:V→UG_B: V \to U to map an image v∈Vv \in V to the domain UU, and an adversarial discriminator DBD_B trained to distinguish generated fake samples GB(v,z′)G_B(v, z') from authentic images u∈Uu \in U.

    The primal and dual tasks form two cyclic reconstruction loops: u→GA(u,z)→GB(GA(u,z),z′)≈uu \to G_A(u, z) \to G_B(G_A(u, z), z') \approx u v→GB(v,z′)→GA(GB(v,z′),z)≈vv \to G_B(v, z') \to G_A(G_B(v, z'), z) \approx v

    This cyclic mapping enables DualGAN to provide two distinct feedback signals to train the translators simultaneously: an adversarial domain membership score evaluated by the discriminators, and a reconstruction error measuring how faithfully the inverted translation recovers the original input image.

  2. Knowl 2 — DualGAN Loss Functions and Objective Formulation

    equation

    DualGAN replaces the standard GAN sigmoid cross-entropy loss with the Wasserstein GAN (WGAN) objective to improve training stability and sample quality, and incorporates an L1L_1 reconstruction loss to enforce cycle consistency without inducing blurriness.

    The discriminator loss functions for DAD_A and DBD_B are formulated as: lAd(u,v)=DA(GA(u,z))−DA(v)l_A^d(u, v) = D_A(G_A(u, z)) - D_A(v) lBd(u,v)=DB(GB(v,z′))−DB(u)l_B^d(u, v) = D_B(G_B(v, z')) - D_B(u)

    where u∈Uu \in U is an image from domain UU, v∈Vv \in V is an image from domain VV, and z,z′z, z' denote latent noise (injected via dropout).

    The generators GAG_A and GBG_B are optimized jointly using a shared generator objective lg(u,v)l^g(u, v): lg(u,v)=λU∥u−GB(GA(u,z),z′)∥1+λV∥v−GA(GB(v,z′),z)∥1−DA(GB(v,z′))−DB(GA(u,z))l^g(u, v) = \lambda_U \|u - G_B(G_A(u, z), z')\|_1 + \lambda_V \|v - G_A(G_B(v, z'), z)\|_1 - D_A(G_B(v, z')) - D_B(G_A(u, z))

    where ∥⋅∥1\|\cdot\|_1 denotes the L1L_1 norm measuring image recovery error, and λU,λV>0\lambda_U, \lambda_V > 0 are constant hyperparameters that weight the reconstruction loss terms relative to the adversarial losses. In practice, λU\lambda_U and λV\lambda_V are selected within the range [100.0,1000.0][100.0, 1000.0]. When domain UU consists of natural photographs and domain VV consists of non-natural representations (such as discrete map data), setting λU<λV\lambda_U < \lambda_V improves output fidelity.

  3. Knowl 3 — DualGAN Training Procedure

    algorithm

    DualGAN is trained using a Wasserstein GAN optimization framework with the RMSProp optimizer, mini-batch stochastic gradient descent, and parameter clipping.

    Input: Unlabeled image sets UU and VV, batch size mm, weight clipping threshold cc, number of discriminator critic steps per generator step ncriticn_{critic}, reconstruction weights λU\lambda_U and λV\lambda_V
    Output: Optimized generator parameters θA,θB\theta_A, \theta_B and discriminator parameters ωA,ωB\omega_A, \omega_B
    Randomly initialize ωA,θA,ωB,θB\omega_A, \theta_A, \omega_B, \theta_B
    repeat
        for t=1t = 1 to ncriticn_{critic} do
            Sample mini-batch of images {u(k)}k=1m⊆U\{u^{(k)}\}_{k=1}^m \subseteq U
            Sample mini-batch of images {v(k)}k=1m⊆V\{v^{(k)}\}_{k=1}^m \subseteq V
            Update ωA\omega_A by descending along its gradient: ∇ωA1m∑k=1m[DA(GA(u(k)))−DA(v(k))]\nabla_{\omega_A} \frac{1}{m} \sum_{k=1}^m [D_A(G_A(u^{(k)})) - D_A(v^{(k)})]
            Update ωB\omega_B by descending along its gradient: ∇ωB1m∑k=1m[DB(GB(v(k)))−DB(u(k))]\nabla_{\omega_B} \frac{1}{m} \sum_{k=1}^m [D_B(G_B(v^{(k)})) - D_B(u^{(k)})]
            ωA←clip(ωA,−c,c)\omega_A \leftarrow \text{clip}(\omega_A, -c, c)
            ωB←clip(ωB,−c,c)\omega_B \leftarrow \text{clip}(\omega_B, -c, c)
        end for
        Sample mini-batch of images {u(k)}k=1m⊆U\{u^{(k)}\}_{k=1}^m \subseteq U
        Sample mini-batch of images {v(k)}k=1m⊆V\{v^{(k)}\}_{k=1}^m \subseteq V
        Update θA,θB\theta_A, \theta_B by descending along their gradients: ∇θA,θB1m∑k=1mlg(u(k),v(k))\nabla_{\theta_A, \theta_B} \frac{1}{m} \sum_{k=1}^m l^g(u^{(k)}, v^{(k)})
    until convergence

    The number of critic steps ncriticn_{critic} is set to 2–42\text{--}4, batch size mm is set to 1–41\text{--}4, and the clipping parameter cc is chosen from [0.01,0.1][0.01, 0.1] depending on the application domain.

  4. Knowl 4 — DualGAN Network Architectures: U-Net and Markovian PatchGAN

    model/method

    DualGAN utilizes symmetric generator and discriminator architectures tailored for high-frequency detail preservation and local structure modeling:

    • Generators (GA,GBG_A, G_B): Both generators use an identical U-Net architecture containing an equal number of downsampling (pooling) and upsampling convolutional layers. Mirrored layers across the contraction and expansion paths are directly connected via skip connections. These skip connections enable low-level visual features (edges, textures, and local shapes) to bypass the bottleneck layer, preventing the loss of high-frequency spatial information during translation. Latent stochasticity is introduced implicitly by applying dropout across several network layers during both training and test phases, rather than concatenating explicit noise vectors to the input.
    • Discriminators (DA,DBD_A, D_B): Both discriminators employ a Markovian PatchGAN architecture with a fixed receptive field of 70×7070 \times 70 pixels. The PatchGAN operates convolutionally across input images (typically of resolution 256×256256 \times 256), modeling local high-frequency textures and styles at the patch level under the assumption of statistical independence beyond the 70×7070 \times 70 spatial patch. All spatial patch responses across the image are averaged to compute the final discriminator score.
  5. Knowl 5 — DualGAN Experimental Setup and Benchmark Tasks

    experimental setup

    DualGAN was evaluated across multiple cross-domain translation tasks spanning supervised benchmark datasets (used in an unpaired manner) and wild web-crawled unpaired collections:

    • Datasets:
      1. PHOTO-SKETCH: Face photos and paired artistic sketches.
      2. DAY-NIGHT: Outdoor scene photos captured during daytime and nighttime.
      3. LABEL-FACADES: Architectural facade photos paired with semantic label maps.
      4. AERIAL-MAPS: Google Map aerial photographs paired with road map graphics.
      5. MATERIAL: Unpaired Flickr photographs of objects categorized into five material classes: stone, metal, plastic, fabric, and wood.
      6. OIL-CHINESE: Unpaired paintings crawled from web search engines consisting of traditional Chinese painting and oil painting styles.
    • Comparative Baselines:
      1. cGAN (pix2pix): Supervised conditional GAN trained on paired images using an L1L_1 or L1+cGANL_1 + \text{cGAN} objective.
      2. Unsupervised GAN: A conditional generator trained with a standard WGAN discriminator by setting λU=λV=0\lambda_U = \lambda_V = 0 in the DualGAN loss, removing cycle reconstruction.
    • Training Details: All models were trained on a single NVIDIA GeForce GTX Titan X GPU at 256×256256 \times 256 resolution. Mini-batch RMSProp was used with batch size m∈[1,4]m \in [1, 4], critic steps ncritic∈[2,4]n_{critic} \in [2, 4], weight clipping parameter c∈[0.01,0.1]c \in [0.01, 0.1], and cycle weights λU,λV∈[100.0,1000.0]\lambda_U, \lambda_V \in [100.0, 1000.0]. Inference latency was under 1 second per image.
  6. Knowl 6 — AMT Realness Scores Across Image Translation Tasks

    data/table

    DualGAN was evaluated via an Amazon Mechanical Turk (AMT) user study assessing perceptual realism against an unsupervised GAN baseline, supervised conditional GAN (cGAN), and authentic ground truth (GT) images. Twenty Turkers scored each synthesized image on a scale from 0 to 4 (0: totally missing, 1: bad, 2: acceptable, 3: good, 4: compelling).

    Task DualGAN cGAN GAN GT
    sketch →\to photo 1.87 1.69 1.04 3.56
    day →\to night 2.42 1.89 0.13 3.05
    label →\to facades 1.89 2.59 1.43 3.33
    map →\to aerial 2.52 2.92 1.88 3.21

    DualGAN outperforms the unsupervised GAN baseline on all four tasks. On tasks where paired training datasets contain spatial misalignments or temporal inconsistencies (such as sketch-to-photo artist variations and day-to-night cloud/object movement), DualGAN outperforms supervised cGAN (1.87 vs. 1.69 on sketch →\to photo; 2.42 vs. 1.89 on day →\to night). On semantic map and facade synthesis tasks, supervised cGAN achieves higher realness scores because pixel-level correspondence provides explicit guidance for coloring semantic labels.

  7. Knowl 7 — Human Perceptual Success Rates in Material Transfer Tasks

    data/table

    The effectiveness of DualGAN for unpaired material transfer was assessed on the MATERIAL dataset using an Amazon Mechanical Turk (AMT) perceptual study. Ten Turkers evaluated 176 translated images across eight directional material transfer tasks. An image translation was deemed a success if at least three out of ten Turkers correctly identified the target material category from the generated image.

    Task DualGAN GAN
    plastic →\to wood 2/11 0/11
    wood →\to plastic 1/11 0/11
    metal →\to stone 2/11 0/11
    stone →\to metal 2/11 0/11
    leather →\to fabric 3/11 2/11
    fabric →\to leather 2/11 1/11
    plastic →\to metal 7/11 3/11
    metal →\to plastic 1/11 0/11

    DualGAN consistently achieves higher material conversion success rates than the baseline GAN across all eight material pairs, demonstrating that the dual reconstruction constraint stabilizes unpaired stylization.

  8. Knowl 8 — Semantic Segmentation Performance on Facade and Aerial Tasks

    data/table

    DualGAN was quantitatively evaluated on semantic parsing tasks (translating natural building facade images to semantic label maps, and aerial photographs to road maps) by computing per-pixel accuracy, per-class accuracy, and class Intersection over Union (Class IoU).

    Method Per-pixel acc. Per-class acc. Class IoU
    Facades →\to Label Task
    DualGAN 0.27 0.13 0.06
    cGAN 0.54 0.33 0.19
    GAN 0.22 0.10 0.05
    Aerial →\to Map Task
    DualGAN 0.42 0.22 0.09
    cGAN 0.70 0.46 0.26
    GAN 0.41 0.23 0.09

    DualGAN outperforms the unsupervised GAN across all segmentation metrics. However, it trails supervised cGAN substantially on both tasks because unsupervised distribution matching cannot uniquely infer precise semantic category assignments without paired ground truth associations.

  9. Knowl 9 — Semantic Misalignment in Unsupervised Image-to-Label Translation

    limitation

    DualGAN relies exclusively on marginal domain distribution matching and cycle reconstruction error without paired pixel-level supervision. While this formulation preserves fine structural contours and local textures, it cannot unambiguously resolve correspondences between continuous visual appearances and discrete semantic label categories.

    As a consequence:

    1. When translating from semantic labels to photographs (such as label →\to facade or map →\to aerial), DualGAN may map valid semantic categories to incorrect textures or colors (e.g., translating orange interstate highway lines into building roofs with bright colors).
    2. When translating from photographs to semantic labels (such as facade →\to label or aerial →\to map), DualGAN frequently outputs sharp, structured geometric segmentations but assigns incorrect discrete label classes to the segmented regions.

Coverage note — No substantial contributed material was omitted. Qualitative visual comparisons across painting stylization and photo-sketch domains were subsumed into the methodological, empirical, and perceptual evaluation knowls.

References

  1. 1.M. Arjovsky, S. Chintala, and L. Bottou. Wasserstein gan. arXiv preprint arXiv:1701.07875, 2017.
  2. 2.Y. Aytar, L. Castrejon, C. Vondrick, H. Pirsiavash, and A. Torralba. Cross-modal scene networks. CoRR, abs/1610.09003, 2016.
  3. 3.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In Advances in neural information processing systems, pages 2672–2680, 2014.
  4. 4.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. arXiv preprint arXiv:1611.07004, 2016.
  5. 5.P.-Y. Laffont, Z. Ren, X. Tao, C. Qian, and J. Hays. Transient attributes for high-level understanding and editing of outdoor scenes. ACM Transactions on Graphics (TOG), 33(4):149, 2014.
  6. 6.A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther. Autoencoding beyond pixels using a learned similarity metric. arXiv preprint arXiv:1512.09300, 2015.
  7. 7.C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, et al. Photo-realistic single image super-resolution using a generative adversarial network. arXiv preprint arXiv:1609.04802, 2016.
  8. 8.C. Li and M. Wand. Precomputed real-time texture synthesis with markovian generative adversarial networks. In European Conference on Computer Vision (ECCV), pages 702–716. Springer, 2016.
  9. 9.M. Liu, T. Breuel, and J. Kautz. Unsupervised image-to-image translation networks. CoRR, abs/1703.00848, 2017.
  10. 10.M.-Y. Liu and O. Tuzel. Coupled generative adversarial networks. In Advances in neural information processing systems, pages 469–477, 2016.
  11. 11.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In Computer Vision and Pattern Recognition (CVPR), pages 3431–3440, 2015.
  12. 12.M. Mathieu, C. Couprie, and Y. LeCun. Deep multi-scale video prediction beyond mean square error. arXiv preprint arXiv:1511.05440, 2015.
  13. 13.M. Mirza and S. Osindero. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  14. 14.G. Perarnau, J. van de Weijer, B. Raducanu, and J. M. Álvarez. Invertible conditional gans for image editing. arXiv preprint arXiv:1611.06355, 2016.
  15. 15.S. Reed, Z. Akata, X. Yan, L. Logeswaran, B. Schiele, and H. Lee. Generative adversarial text to image synthesis. In Proceedings of The 33rd International Conference on Machine Learning, volume 3, 2016.
  16. 16.O. Ronneberger, P. Fischer, and T. Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 234–241. Springer, 2015.
  17. 17.L. Sharan, R. Rosenholtz, and E. Adelson. Material perception: What can you see in a brief glance? Journal of Vision, 9(8):784–784, 2009.
  18. 18.Y. Taigman, A. Polyak, and L. Wolf. Unsupervised cross-domain image generation. arXiv preprint arXiv:1611.02200, 2016.
  19. 19.T. Tieleman and G. Hinton. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning, 4(2), 2012.
  20. 20.R. Tyleček and R. Šára. Spatial pattern templates for recognition of objects with regular structure. In German Conference on Pattern Recognition, pages 364–374. Springer, 2013.
  21. 21.X. Wang and A. Gupta. Generative image modeling using style and structure adversarial networks. In European Conference on Computer Vision (ECCV), pages 318–335. Springer, 2016.
  22. 22.X. Wang and X. Tang. Face photo-sketch synthesis and recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 31(11):1955–1967, 2009.
  23. 23.Y. Xia, D. He, T. Qin, L. Wang, N. Yu, T.-Y. Liu, and W.-Y. Ma. Dual learning for machine translation. arXiv preprint arXiv:1611.00179, 2016.
  24. 24.X. Yan, J. Yang, K. Sohn, and H. Lee. Attribute2image: Conditional image generation from visual attributes. In European Conference on Computer Vision (ECCV), pages 776–791. Springer, 2016.
  25. 25.W. Zhang, X. Wang, and X. Tang. Coupled information-theoretic encoding for face photo-sketch recognition. In Computer Vision and Pattern Recognition (CVPR), pages 513–520. IEEE, 2011.
  26. 26.J. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In International Conference on Computer Vision (ICCV), to appear, 2017.

Citation

MLA
Yi, Z., et al. “DualGAN: Unsupervised Dual Learning for Image-to-Image Translation”. arXiv, 2017, http://arxiv.org/abs/1704.02510v4.
APA
Yi, Z., Zhang, H., Tan, P., & Gong, M. (2017). DualGAN: Unsupervised Dual Learning for Image-to-Image Translation. arXiv. http://arxiv.org/abs/1704.02510v4
Chicago
Yi, Z., H. Zhang, P. Tan, and M. Gong. 2017. “DualGAN: Unsupervised Dual Learning for Image-to-Image Translation”. arXiv. http://arxiv.org/abs/1704.02510v4.
Harvard
Yi, Z. et al. (2017) “DualGAN: Unsupervised Dual Learning for Image-to-Image Translation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1704.02510v4.
Vancouver
1. Yi Z, Zhang H, Tan P, Gong M (2017) DualGAN: Unsupervised Dual Learning for Image-to-Image Translation. arXiv

BibTeX

@article{yi2017dualgan,
  title = {DualGAN: Unsupervised Dual Learning for Image-to-Image Translation},
  author = {Yi, Zili and Zhang, Hao and Tan, Ping and Gong, Minglun},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1704.02510v4},
  eprint = {1704.02510}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE