ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation
Tuan-Hung VuHimalaya JainMaxime BucherMatthieu CordPatrick Pérez
Introduces an adversarial entropy minimization framework that significantly narrows the synthetic-to-real performance gap in semantic segmentation by driving models to produce confident and structurally consistent predictions on unlabeled target data.
Deploying artificial intelligence systems in safety-critical applications, such as autonomous driving, requires deep neural networks to maintain high visual recognition accuracy across diverse environments. While training models on inexpensive synthetic data from video game engines offers massive cost and time advantages over manual pixel-level labeling, models typically suffer severe performance loss when transferred to real-world conditions due to the distribution shift between source and target domains.
The article demonstrates an unsupervised domain adaptation framework based on prediction uncertainty minimization to adapt semantic segmentation models trained on labeled synthetic data to unlabeled real-world target environments.
The authors evaluate two complementary approaches: direct entropy minimization, which directly penalizes pixel-level uncertainty with negligible computational overhead, and adversarial entropy minimization, which aligns prediction certainty and spatial layout structure across domains using a discriminator network. Testing was conducted across two standard synthetic-to-real benchmarks (GTA5 to Cityscapes and SYNTHIA to Cityscapes) using 2,975 unlabeled training images and standard convolutional backbones, with an additional extension to object detection under adverse weather conditions.
The evaluation produced four key findings. First, adversarial entropy minimization established state-of-the-art segmentation accuracy, achieving a 43.8% mean intersection-over-union on GTA5 to Cityscapes and 47.6% on SYNTHIA to Cityscapes with a ResNet backbone. Second, ensembling direct and adversarial entropy minimization yielded the highest performance, reaching 45.5% and 48.0% respectively, and significantly reduced the accuracy gap relative to fully supervised models. Third, direct entropy minimization achieved competitive baseline adaptation with minimal training overhead and greater numerical stability than typical adversarial methods. Fourth, when applied to object detection in foggy environments, the adversarial framework improved mean average precision from a 14.7% baseline to 26.2%, an absolute gain of 11.5%.
These findings indicate that enforcing confident, low-entropy predictions on unlabeled target data successfully bridges the domain gap while avoiding expensive target annotations. The direct entropy approach provides an operationally stable, low-cost adaptation technique, whereas the adversarial approach effectively captures complex spatial structures. However, unconstrained entropy minimization can risk biasing models toward dominant or easy classes when domain layouts differ substantially, which the article mitigates by incorporating relaxed source class-ratio priors.
Organizations developing computer vision systems should adopt entropy-based adaptation to lower annotation costs and improve generalization. Teams prioritizing compute efficiency and stable training should implement direct entropy minimization, while applications demanding maximum accuracy should employ the adversarial model or an ensemble of both. For production pipelines with large structural shifts, practitioners should integrate class-ratio priors. Future work should expand these entropy techniques to larger, modern object detection architectures and evaluate performance across a wider variety of operational weather and lighting conditions.
- Paper: Learning to Adapt Structured Output Space for Semantic Segmentation, Yi-Hsuan Tsai et al. (2018). It introduced adversarial output-space adaptation for semantic segmentation, establishing the direct architectural and loss formulation foundation that ADVENT builds upon and enhances via entropy minimization.
- Paper: CyCADA: Cycle-Consistent Adversarial Domain Adaptation, Judy Hoffman et al. (2018). It formalizes cycle-consistent and adversarial domain adaptation across pixel and feature spaces for synthetic-to-real segmentation benchmarks, which ADVENT adopts as core evaluation settings.
- Book: Domain-Adversarial Training of Neural Networks, Yaroslav Ganin et al. (2016). It establishes the foundational domain-adversarial framework and gradient reversal concepts that underpin adversarial domain alignment architectures in deep networks.
- Paper: Adversarial Discriminative Domain Adaptation, Eric Tzeng et al. (2017). It provides a generalized adversarial framework for unsupervised discriminative domain adaptation that ADVENT refines for dense prediction tasks.
- Paper: Conditional Adversarial Domain Adaptation, Mingsheng Long et al. (2017). It explores conditioning adversarial alignment on prediction uncertainty and entropy, motivating ADVENT's direct minimization of pixel-wise prediction entropy.
- Paper: Maximum Classifier Discrepancy for Unsupervised Domain Adaptation, Kuniaki Saito et al. (2017). It demonstrates aligning representations near task decision boundaries in unsupervised domain adaptation, providing context for confidence- and boundary-aware adaptation.
- Paper: The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes, Germán Ros et al. (2016). It introduces the SYNTHIA benchmark for urban semantic segmentation, serving as one of the primary synthetic-to-real datasets used to evaluate ADVENT.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). It defines the end-to-end fully convolutional network paradigm for semantic segmentation on which modern dense prediction and adaptation models rely.
- Paper: Rethinking Atrous Convolution for Semantic Image Segmentation, Liang-Chieh Chen et al. (2017). It develops the DeepLabv3 atrous spatial pyramid pooling architecture that serves as the backbone segmentation network in domain adaptation frameworks.
- Paper: Tent: Fully Test-Time Adaptation by Entropy Minimization, Dequan Wang et al. (2021). It extends the concept of entropy minimization to the fully test-time adaptation setting by directly optimizing model parameters on unlabeled target test data.
- Paper: Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation, Jian Liang et al. (2020). It advances unsupervised adaptation to the source-free setting using information maximization and pseudo-labeling, building on confidence-driven objectives like entropy minimization.
- Paper: Deep Visual Domain Adaptation: A Survey, Mei Wang et al. (2018). It provides a comprehensive survey classifying adversarial, discrepancy, and reconstruction paradigms in visual domain adaptation, including methods developed in ADVENT.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). It surveys strategies to generalize to unseen target environments without requiring target data access during training, expanding beyond unsupervised domain adaptation.
- Paper: WILDS: A Benchmark of in-the-Wild Distribution Shifts, Pang Wei Koh et al. (2020). It broadens the evaluation of distribution shifts to diverse real-world benchmarks across multiple modalities, testing adaptation and robustness strategies in realistic settings.
