ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation

Tuan-Hung VuHimalaya JainMaxime BucherMatthieu CordPatrick Pérez

article2018CVPR1,549 citations

Introduces an adversarial entropy minimization framework that significantly narrows the synthetic-to-real performance gap in semantic segmentation by driving models to produce confident and structurally consistent predictions on unlabeled target data.

Listen

Deploying artificial intelligence systems in safety-critical applications, such as autonomous driving, requires deep neural networks to maintain high visual recognition accuracy across diverse environments. While training models on inexpensive synthetic data from video game engines offers massive cost and time advantages over manual pixel-level labeling, models typically suffer severe performance loss when transferred to real-world conditions due to the distribution shift between source and target domains.

The article demonstrates an unsupervised domain adaptation framework based on prediction uncertainty minimization to adapt semantic segmentation models trained on labeled synthetic data to unlabeled real-world target environments.

The authors evaluate two complementary approaches: direct entropy minimization, which directly penalizes pixel-level uncertainty with negligible computational overhead, and adversarial entropy minimization, which aligns prediction certainty and spatial layout structure across domains using a discriminator network. Testing was conducted across two standard synthetic-to-real benchmarks (GTA5 to Cityscapes and SYNTHIA to Cityscapes) using 2,975 unlabeled training images and standard convolutional backbones, with an additional extension to object detection under adverse weather conditions.

The evaluation produced four key findings. First, adversarial entropy minimization established state-of-the-art segmentation accuracy, achieving a 43.8% mean intersection-over-union on GTA5 to Cityscapes and 47.6% on SYNTHIA to Cityscapes with a ResNet backbone. Second, ensembling direct and adversarial entropy minimization yielded the highest performance, reaching 45.5% and 48.0% respectively, and significantly reduced the accuracy gap relative to fully supervised models. Third, direct entropy minimization achieved competitive baseline adaptation with minimal training overhead and greater numerical stability than typical adversarial methods. Fourth, when applied to object detection in foggy environments, the adversarial framework improved mean average precision from a 14.7% baseline to 26.2%, an absolute gain of 11.5%.

These findings indicate that enforcing confident, low-entropy predictions on unlabeled target data successfully bridges the domain gap while avoiding expensive target annotations. The direct entropy approach provides an operationally stable, low-cost adaptation technique, whereas the adversarial approach effectively captures complex spatial structures. However, unconstrained entropy minimization can risk biasing models toward dominant or easy classes when domain layouts differ substantially, which the article mitigates by incorporating relaxed source class-ratio priors.

Organizations developing computer vision systems should adopt entropy-based adaptation to lower annotation costs and improve generalization. Teams prioritizing compute efficiency and stable training should implement direct entropy minimization, while applications demanding maximum accuracy should employ the adversarial model or an ensemble of both. For production pipelines with large structural shifts, practitioners should integrate class-ratio priors. Future work should expand these entropy techniques to larger, modern object detection architectures and evaluate performance across a wider variety of operational weather and lighting conditions.

Cover for ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation

Abstract

Semantic segmentation is a key problem for many computer vision tasks. While approaches based on convolutional neural networks constantly break new records on different benchmarks, generalizing well to diverse testing environments remains a major challenge. In numerous real world applications, there is indeed a large gap between data distributions in train and test domains, which results in severe performance loss at run-time. In this work, we address the task of unsupervised domain adaptation in semantic segmentation with losses based on the entropy of the pixel-wise predictions. To this end, we propose two novel, complementary methods using (i) entropy loss and (ii) adversarial loss respectively. We demonstrate state-of-the-art performance in semantic segmentation on two challenging "synthetic-2-real" set-ups and show that the approach can also be used for detection.

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 3 Approaches
  • 3.1 Direct entropy minimization
  • 3.2 Minimizing entropy with adversarial learning
  • 3.3 Incorporating class-ratio priors
  • 4 Experiments
  • 4.1 Experimental details
  • 4.2 Results
  • 4.3 Discussion
  • 5 Conclusion
  • A Entropy-based UDA for object detection
  • References

Knowls

  1. Knowl 1 — Direct Entropy Minimization for Semantic Segmentation Domain Adaptation

    model/method

    Direct entropy minimization (MinEnt) addresses unsupervised domain adaptation in semantic segmentation by maximizing prediction certainty directly on unlabeled target images. Let Xs⊂RH×W×3X_s \subset \mathbb{R}^{H \times W \times 3} denote the labeled source dataset with corresponding ground-truth segmentation masks Ys⊂{1,…,C}H×WY_s \subset \{1, \dots, C\}^{H \times W}, where CC is the number of semantic classes and each pixel label ys(h,w)y_s^{(h,w)} is a one-hot vector ys(h,w,c)∈{0,1}y_s^{(h,w,c)} \in \{0, 1\}. Let Xt⊂RH×W×3X_t \subset \mathbb{R}^{H \times W \times 3} be the unlabeled target dataset.

    A segmentation network FF with parameters θF\theta_F predicts class probability maps Px=F(x)∈[0,1]H×W×CP_x = F(x) \in [0, 1]^{H \times W \times C}, where Px(h,w,c)P_x^{(h,w,c)} is the predicted probability for class cc at pixel location (h,w)(h, w). For a target image xt∈Xtx_t \in X_t, the normalized Shannon entropy Ext(h,w)∈[0,1]E_{x_t}^{(h,w)} \in [0, 1] at pixel (h,w)(h, w) is given by:

    Ext(h,w)=−1log⁡(C)∑c=1CPxt(h,w,c)log⁡Pxt(h,w,c)E_{x_t}^{(h,w)} = \frac{-1}{\log(C)} \sum_{c=1}^C P_{x_t}^{(h,w,c)} \log P_{x_t}^{(h,w,c)}

    The total target entropy loss Lent(xt)L_{ent}(x_t) is computed by summing the normalized pixel-wise entropies over all spatial locations:

    Lent(xt)=∑h=1H∑w=1WExt(h,w)L_{ent}(x_t) = \sum_{h=1}^H \sum_{w=1}^W E_{x_t}^{(h,w)}

    The network parameters θF\theta_F are optimized jointly on source cross-entropy segmentation loss LsegL_{seg} and target unsupervised entropy loss LentL_{ent}:

    min⁡θF1∣Xs∣∑xs∈XsLseg(xs,ys)+λent∣Xt∣∑xt∈XtLent(xt)\min_{\theta_F} \frac{1}{|X_s|} \sum_{x_s \in X_s} L_{seg}(x_s, y_s) + \frac{\lambda_{ent}}{|X_t|} \sum_{x_t \in X_t} L_{ent}(x_t)

    where Lseg(xs,ys)=−∑h=1H∑w=1W∑c=1Cys(h,w,c)log⁡Pxs(h,w,c)L_{seg}(x_s, y_s) = -\sum_{h=1}^H \sum_{w=1}^W \sum_{c=1}^C y_s^{(h,w,c)} \log P_{x_s}^{(h,w,c)} and λent>0\lambda_{ent} > 0 is a loss weighting hyperparameter (set to λent=0.001\lambda_{ent} = 0.001).

  2. Knowl 2 — Adversarial Entropy Minimization via Weighted Self-Information Alignment

    model/method

    Adversarial entropy minimization (AdvEnt) minimizes prediction entropy indirectly by matching the distribution of weighted self-information between source and target predictions. This alignment leverages structural scene layout dependencies across pixels, which are lost when pixel entropies are minimized independently.

    For an input image xx, the segmentation network FF produces class probabilities Px∈[0,1]H×W×CP_x \in [0, 1]^{H \times W \times C}. The weighted self-information map Ix∈RH×W×CI_x \in \mathbb{R}^{H \times W \times C} is defined at pixel (h,w)(h, w) and class cc as:

    Ix(h,w,c)=−Px(h,w,c)log⁡Px(h,w,c)I_x^{(h,w,c)} = - P_x^{(h,w,c)} \log P_x^{(h,w,c)}

    Summing Ix(h,w,c)I_x^{(h,w,c)} across all classes cc recovers the unnormalized Shannon entropy at that pixel.

    A fully-convolutional discriminator DD with parameters θD\theta_D takes IxI_x as input and produces domain classification predictions, where source corresponds to label 1 and target corresponds to label 0. The discriminator objective function is:

    min⁡θD1∣Xs∣∑xs∈XsLD(Ixs,1)+1∣Xt∣∑xt∈XtLD(Ixt,0)\min_{\theta_D} \frac{1}{|X_s|} \sum_{x_s \in X_s} L_D(I_{x_s}, 1) + \frac{1}{|X_t|} \sum_{x_t \in X_t} L_D(I_{x_t}, 0)

    where LDL_D is the binary cross-entropy domain classification loss. The segmentation network FF is trained to maximize the domain confusion of target self-information maps while minimizing supervised source cross-entropy:

    min⁡θF1∣Xs∣∑xs∈XsLseg(xs,ys)+λadv∣Xt∣∑xt∈XtLD(Ixt,1)\min_{\theta_F} \frac{1}{|X_s|} \sum_{x_s \in X_s} L_{seg}(x_s, y_s) + \frac{\lambda_{adv}}{|X_t|} \sum_{x_t \in X_t} L_D(I_{x_t}, 1)

    with weighting factor λadv=0.001\lambda_{adv} = 0.001. Because models trained with source supervision naturally output low-entropy predictions on source images, matching the target self-information distribution to the source distribution implicitly minimizes target prediction entropy while transferring structural consistency.

  3. Knowl 3 — Relationship Between Entropy Minimization and Self-Training Pseudo-Labeling

    theoretical result

    In self-training (ST) for unsupervised domain adaptation, pseudo-labels y^t∈{0,1}H×W×C\hat{y}_t \in \{0, 1\}^{H \times W \times C} are generated for a subset of confident pixels K⊂{1,…,H}×{1,…,W}K \subset \{1, \dots, H\} \times \{1, \dots, W\} where prediction confidence exceeds a given threshold, and the model minimizes:

    Lseg(xt,y^t)=−∑(h,w)∈K∑c=1Cy^t(h,w,c)log⁡Pxt(h,w,c)L_{seg}(x_t, \hat{y}_t) = -\sum_{(h,w) \in K} \sum_{c=1}^C \hat{y}_t^{(h,w,c)} \log P_{x_t}^{(h,w,c)}

    Direct entropy minimization optimizes the normalized Shannon entropy across all spatial coordinates without hard thresholding:

    Lent(xt)=−1log⁡(C)∑h=1H∑w=1W∑c=1CPxt(h,w,c)log⁡Pxt(h,w,c)L_{ent}(x_t) = \frac{-1}{\log(C)} \sum_{h=1}^H \sum_{w=1}^W \sum_{c=1}^C P_{x_t}^{(h,w,c)} \log P_{x_t}^{(h,w,c)}

    Direct entropy minimization is equivalent to a soft-assignment formulation of the pseudo-label cross-entropy loss. Unlike self-training, which relies on a discrete selection set KK and threshold scheduling, entropy minimization continuously weights the class log-likelihoods by their own predicted probabilities Pxt(h,w,c)P_{x_t}^{(h,w,c)} across all pixels.

  4. Knowl 4 — Class-Ratio Prior Regularization for Entropy Minimization

    model/method

    When large domain discrepancies in camera viewpoint and layout exist between source and target datasets, entropy minimization can disproportionately favor dominant classes and suppress rare categories. Class-ratio prior regularization enforces class diversity on target predictions by aligning target class frequencies with source label statistics.

    Let ps∈[0,1]Cp_s \in [0, 1]^C denote the ℓ1\ell_1-normalized class histogram computed over all annotated pixels across the source training set. For a target image xtx_t, the spatial expectation of the predicted probability for class cc is defined as Ec(Pxt(c))=1H⋅W∑h=1H∑w=1WPxt(h,w,c)\mathbb{E}_c(P_{x_t}^{(c)}) = \frac{1}{H \cdot W} \sum_{h=1}^H \sum_{w=1}^W P_{x_t}^{(h,w,c)}. The class-prior loss Lcp(xt)L_{cp}(x_t) penalizes predictions where the expected class probability falls below a relaxed fraction of the source class proportion:

    Lcp(xt)=∑c=1Cmax⁡(0, μ ps(c)−Ec(Pxt(c)))L_{cp}(x_t) = \sum_{c=1}^C \max\left(0, \, \mu \, p_s^{(c)} - \mathbb{E}_c(P_{x_t}^{(c)})\right)

    where μ∈[0,1]\mu \in [0, 1] is a relaxation parameter. Setting μ=0\mu = 0 imposes no constraint, while μ=1\mu = 1 forces single-image target distributions to strictly mirror the source domain ratio. Setting μ=0.5\mu = 0.5 balances the preservation of minority classes with the inherent variations across individual target images.

  5. Knowl 5 — Target Entropy Minimization on Confused Pixel Ranges

    model/method

    For high-capacity segmentation networks (such as ResNet-101), direct entropy minimization can be made more effective by applying the loss exclusively to target pixels with high uncertainty (entropy range selection, MinEnt+ER).

    For each target image xtx_t, normalized pixel-wise entropies Ext(h,w)E_{x_t}^{(h,w)} are computed, and the entropy loss is calculated only over the top 30% highest-entropy spatial locations within that image:

    LentER(xt)=∑(h,w)∈Ωtop30(xt)Ext(h,w)L_{ent}^{ER}(x_t) = \sum_{(h,w) \in \Omega_{top30}(x_t)} E_{x_t}^{(h,w)}

    where Ωtop30(xt)\Omega_{top30}(x_t) denotes the set containing the 30% of image pixels having the highest entropy values. Because high-capacity models produce predictions on ambiguous regions that are frequently accurate but under-confident, restricting minimization to these indecisive pixels drives the decision boundaries toward low-density regions without reinforcing erroneous high-confidence predictions.

  6. Knowl 6 — Semantic Segmentation Adaptation Performance on GTA5 to Cityscapes

    empirical result

    On the synthetic-to-real domain adaptation benchmark from GTA5 (24,966 synthetic training images, 19 shared classes) to Cityscapes (2,975 unlabeled training images, 500 validation images), models trained using direct entropy minimization (MinEnt), adversarial entropy minimization (AdvEnt), and their ensemble achieve state-of-the-art mean Intersection-over-Union (mIoU):

    Backbone Adaptation Method mIoU (%) Oracle Gap (%)
    VGG-16 FCNs in the Wild 27.1 -37.5
    VGG-16 CyCADA 34.8 -31.4
    VGG-16 Adapt-SegMap 35.0 -26.8
    VGG-16 Self-Training + Class Balancing 30.9 -
    VGG-16 Ours (MinEnt) 32.8 -
    VGG-16 Ours (AdvEnt) 36.1 -25.7
    ResNet-101 Adapt-SegMap 42.4 -22.7
    ResNet-101 Adapt-SegMap* (Retrained) 42.2 -
    ResNet-101 Ours (MinEnt) 42.3 -
    ResNet-101 Ours (MinEnt + ER) 43.1 -
    ResNet-101 Ours (AdvEnt) 43.8 -21.3
    ResNet-101 Ours (AdvEnt + MinEnt Ensemble) 45.5 -19.6

    The oracle ResNet-101 model trained with full target supervision reaches 65.1% mIoU. The ensemble of AdvEnt and MinEnt achieves 45.5% mIoU, reducing the oracle performance gap to 19.6%.

  7. Knowl 7 — Semantic Segmentation Adaptation Performance on SYNTHIA to Cityscapes

    empirical result

    On the synthetic-to-real adaptation benchmark from SYNTHIA-RAND-CITYSCAPES (9,400 synthetic images) to Cityscapes (evaluated over all 16 shared classes and a standard 13-class subset), entropy minimization combined with class-ratio priors (CP) achieves state-of-the-art segmentation performance:

    Backbone Adaptation Method 16-class mIoU (%) 13-class mIoU* (%)
    VGG-16 FCNs in the Wild 20.2 22.1
    VGG-16 Adapt-SegMap - 37.6
    VGG-16 Self-Training 23.9 27.8
    VGG-16 Self-Training + Class Balancing 35.4 36.1
    VGG-16 Ours (MinEnt) 27.5 32.5
    VGG-16 Ours (MinEnt + CP) 30.4 35.4
    VGG-16 Ours (AdvEnt + CP) 31.4 36.6
    ResNet-101 Adapt-SegMap - 46.7
    ResNet-101 Adapt-SegMap* (Retrained) 39.6 45.8
    ResNet-101 Ours (MinEnt) 38.1 44.2
    ResNet-101 Ours (AdvEnt) 40.8 47.6
    ResNet-101 Ours (AdvEnt + MinEnt Ensemble) 41.2 48.0

    Adding the class-prior regularization (+CP) improves VGG-16 MinEnt by +2.9% mIoU on both 16- and 13-class subsets. On ResNet-101, the ensemble of AdvEnt and MinEnt achieves 41.2% (16 classes) and 48.0% (13 classes), exceeding prior structured adaptation baselines.

  8. Knowl 8 — Extension of Entropy Minimization to Unsupervised Domain Adaptation in Object Detection

    model/method

    Entropy minimization principles generalize to anchor-based multi-scale object detectors, such as the Single Shot MultiBox Detector (SSD-300). For an input image xx, SSD-300 predicts dense bounding boxes across MM feature maps at distinct spatial resolutions. At feature map m∈{1,…,M}m \in \{1, \dots, M\} of dimensions Hm×WmH_m \times W_m with KmK_m anchor boxes per cell, the predicted soft-detection distribution over CC classes is Pxm∈[0,1]Hm×Wm×Km×CP_x^m \in [0, 1]^{H_m \times W_m \times K_m \times C}.

    1. Direct Entropy Minimization for Detection: The normalized box-level Shannon entropy for anchor box kk at cell (h,w)(h, w) on feature map mm is:

    Extm(h,w,k)=−1log⁡(C)∑c=1CPxtm(h,w,k,c)log⁡Pxtm(h,w,k,c)E_{x_t}^{m(h,w,k)} = \frac{-1}{\log(C)} \sum_{c=1}^C P_{x_t}^{m(h,w,k,c)} \log P_{x_t}^{m(h,w,k,c)}

    The total detection entropy loss Lent(xt)L_{ent}(x_t) sums over all anchor boxes, cells, and feature resolutions:

    Lent(xt)=∑m=1M∑h=1Hm∑w=1Wm∑k=1KmExtm(h,w,k)L_{ent}(x_t) = \sum_{m=1}^M \sum_{h=1}^{H_m} \sum_{w=1}^{W_m} \sum_{k=1}^{K_m} E_{x_t}^{m(h,w,k)}

    1. Adversarial Entropy Minimization for Detection: For each feature map mm, weighted self-information tensors IxmI_x^m are formed with entries −Pxm(h,w,k,c)log⁡Pxm(h,w,k,c)-P_x^{m(h,w,k,c)} \log P_x^{m(h,w,k,c)}. The maps IxmI_x^m at smaller spatial resolutions are zero-padded to match the spatial size of the largest feature map (H1×W1H_1 \times W_1) and concatenated along the channel dimension into a tensor IxI_x, which is then fed into a convolutional domain discriminator.
  9. Knowl 9 — Object Detection Domain Adaptation Performance on Cityscapes to Cityscapes Foggy

    empirical result

    When evaluated on the Cityscapes →\to Cityscapes Foggy domain adaptation setup using an SSD-300 detector (VGG-16 backbone) at 300×300300 \times 300 resolution, entropy minimization methods improve detection mean average precision (mAP) over the unadapted source-only baseline across all 8 annotated classes:

    Model person rider car truck bus train mcycle bicycle mAP (%)
    SSD-300 (Source Baseline) 15.0 17.4 27.2 5.7 15.1 9.1 11.0 16.7 14.7
    SSD-300 + MinEnt 15.8 22.0 28.3 5.0 15.2 15.0 13.0 20.6 16.9
    SSD-300 + AdvEnt 17.6 25.0 39.6 20.0 37.1 25.9 21.3 23.1 26.2

    Adversarial entropy adaptation (AdvEnt) improves overall mAP from 14.7% to 26.2% (+11.5% mAP), with significant increases on classes such as car (27.2% to 39.6%), truck (5.7% to 20.0%), and bus (15.1% to 37.1%).

Coverage note — None was omitted; all contributed models, equations, extensions (ER, CP, detection), and experimental results were included.

References

  1. 1.L. Bottou. Large-scale machine learning with stochastic gradient descent. In COMPSTAT. 2010. 5
  2. 2.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. PAMI, 2018. 1, 5
  3. 3.Y. Chen, W. Li, C. Sakaridis, D. Dai, and L. Van Gool. Domain adaptive faster r-cnn for object detection in the wild. In CVPR, 2018. 8
  4. 4.Y. Chen, W. Li, and L. Van Gool. Road: Reality oriented adaptation for semantic segmentation of urban scenes. In CVPR, 2018. 2
  5. 5.Y.-H. Chen, W.-Y. Chen, Y.-T. Chen, B.-C. Tsai, Y.-C. F. Wang, and M. Sun. No more discrimination: Cross city adaptation of road scene segmenters. In ICCV, 2017. 2
  6. 6.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. 5
  7. 7.G. Csurka. Domain adaptation for visual applications: A comprehensive survey. 2017. 2
  8. 8.A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun. CARLA: An open urban driving simulator. In CoRL, 2017. 2
  9. 9.M. Everingham, S. M. A. Eslami, C. K. I. Williams, J. Winn, and A. Zisserman. The pascal visual object classes challenge: A retrospective. IJCV, 2015. 5
  10. 10.Y. Ganin and V. Lempitsky. Unsupervised domain adaptation by backpropagation. In ICML, 2015. 1, 2
  11. 11.I. Goodfellow, Y. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio. Generative adversarial nets. In NIPS, 2014. 3, 4
  12. 12.Y. Grandvalet and Y. Bengio. Semi-supervised learning by entropy minimization. In NIPS, 2005. 3
  13. 13.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In CVPR, 2016. 5
  14. 14.J. Hoffman, E. Tzeng, T. Park, J.-Y. Zhu, P. Isola, K. Saenko, A. Efros, and T. Darrell. CyCADA: Cycle-consistent adversarial domain adaptation. In ICML, 2018. 1, 2, 5, 6, 7
  15. 15.J. Hoffman, D. Wang, F. Yu, and T. Darrell. Fcns in the wild: Pixel-level adversarial and constraint-based adaptation. arXiv preprint arXiv:1612.02649, 2016. 1, 2, 5, 6, 7
  16. 16.W. Hong, Z. Wang, M. Yang, and J. Yuan. Conditional generative adversarial network for structured domain adaptation. In CVPR, 2018. 2
  17. 17.H. Jain, J. Zepeda, P. Perez, and R. Gribonval. Subic: A supervised, structured binary code for image search. In ICCV, 2017. 3
  18. 18.H. Jain, J. Zepeda, P. Perez, and R. Gribonval. Learning a complete image indexing pipeline. In CVPR, 2018. 3
  19. 19.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 5
  20. 20.S. Laine and T. Aila. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242, 2016. 3
  21. 21.D.-H. Lee. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In ICML Workshop, 2013. 3
  22. 22.W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. E. Reed, C. Fu, and A. C. Berg. SSD: single shot multibox detector. In ECCV, 2016. 8
  23. 23.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 1
  24. 24.M. Long, Y. Cao, J. Wang, and M. I. Jordan. Learning transferable features with deep adaptation networks. In ICML, 2015. 1, 2
  25. 25.M. Long, H. Zhu, J. Wang, and M. I. Jordan. Unsupervised domain adaptation with residual transfer networks. In NIPS, 2016. 1, 2, 3
  26. 26.Z. Murez, S. Kolouri, D. Kriegman, R. Ramamoorthi, and K. Kim. Image to image translation for domain adaptation. In CVPR, 2018. 3
  27. 27.A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer. Automatic differentiation in pytorch. In NIPS Workshop, 2017. 5
  28. 28.A. Radford, L. Metz, and S. Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks. In ICLR, 2016. 5
  29. 29.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NIPS, 2015. 8
  30. 30.S. R. Richter, V. Vineet, S. Roth, and V. Koltun. Playing for data: Ground truth from computer games. In ECCV, 2016. 1, 2, 5
  31. 31.G. Ros, L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez. The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In CVPR, 2016. 1, 2, 5
  32. 32.K. Saito, Y. Ushiku, T. Harada, and K. Saenko. Adversarial dropout regularization. In ICLR, 2018. 1, 2
  33. 33.K. Saito, K. Watanabe, Y. Ushiku, and T. Harada. Maximum classifier discrepancy for unsupervised domain adaptation. In CVPR, 2018. 2
  34. 34.S. Sankaranarayanan, Y. Balaji, A. Jain, S. Nam Lim, and R. Chellappa. Learning from synthetic data: Addressing domain shift for semantic segmentation. In CVPR, 2018. 1, 2
  35. 35.A. Shafaei, J. J. Little, and M. Schmidt. Play and learn: Using video games to train computer vision models. In BMVC, 2016. 2
  36. 36.C. E. Shannon. A mathematical theory of communication. Bell system technical journal, 1948. 3
  37. 37.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015. 5, 8
  38. 38.J. T. Springenberg. Unsupervised and semi-supervised learning with categorical generative adversarial networks. ICLR, 2016. 2, 3
  39. 39.A. Tarvainen and H. Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In NIPS, 2017. 3
  40. 40.M. Tribus. Thermostatics and thermodynamics. 1970. 4
  41. 41.Y.-H. Tsai, W.-C. Hung, S. Schulter, K. Sohn, M.-H. Yang, and M. Chandraker. Learning to adapt structured output space for semantic segmentation. In CVPR, 2018. 2, 4, 5, 6, 7
  42. 42.E. Tzeng, J. Hoffman, K. Saenko, and T. Darrell. Adversarial discriminative domain adaptation. In CVPR, 2017. 1, 2
  43. 43.Z. Wu, X. Han, Y.-L. Lin, M. Gokhan Uzunbas, T. Goldstein, S. Nam Lim, and L. S. Davis. Dcan: Dual channel-wise alignment networks for unsupervised scene adaptation. In ECCV, 2018. 1, 2, 3
  44. 44.H. Yan, Y. Ding, P. Li, Q. Wang, Y. Xu, and W. Zuo. Mind the class weight bias: Weighted maximum mean discrepancy for unsupervised domain adaptation. In CVPR, 2017. 1
  45. 45.Y. Zhang, P. David, and B. Gong. Curriculum domain adaptation for semantic segmentation of urban scenes. In ICCV, 2017. 3
  46. 46.Y. Zhang, Z. Qiu, T. Yao, D. Liu, and T. Mei. Fully convolutional adaptation networks for semantic segmentation. In CVPR, 2018. 3
  47. 47.H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia. Pyramid scene parsing network. In CVPR, 2017. 1
  48. 48.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV, 2017. 2
  49. 49.X. Zhu. Semi-supervised learning literature survey. Technical Report 1530, Computer Sciences, University of Wisconsin-Madison, 2005. 2
  50. 50.X. Zhu, H. Zhou, C. Yang, J. Shi, and D. Lin. Penalizing top performers: Conservative loss for semantic segmentation adaptation. In ECCV, September 2018. 3
  51. 51.Y. Zou, Z. Yu, B. V. Kumar, and J. Wang. Unsupervised domain adaptation for semantic segmentation via class-balanced self-training. In ECCV, 2018. 1, 2, 3, 4, 5, 6, 7

Citation

MLA
Vu, T.-H., et al. “ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation”. arXiv, 2018, http://arxiv.org/abs/1811.12833v2.
APA
Vu, T.-H., Jain, H., Bucher, M., Cord, M., & Pérez, P. (2018). ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation. arXiv. http://arxiv.org/abs/1811.12833v2
Chicago
Vu, T.-H., H. Jain, M. Bucher, M. Cord, and P. Pérez. 2018. “ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation”. arXiv. http://arxiv.org/abs/1811.12833v2.
Harvard
Vu, T.-H. et al. (2018) “ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1811.12833v2.
Vancouver
1. Vu T-H, Jain H, Bucher M, Cord M, Pérez P (2018) ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation. arXiv

BibTeX

@article{vu2018advent,
  title = {ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation},
  author = {Vu, Tuan-Hung and Jain, Himalaya and Bucher, Maxime and Cord, Matthieu and Pérez, Patrick},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1811.12833v2},
  eprint = {1811.12833}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE