Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training

Yang ZouZhiding YuB. V. K. Vijaya KumarJinsong Wang

article2018ECCV1,518 citations

Proposes a class-balanced self-training framework for unsupervised domain adaptation in semantic segmentation that prevents dominant categories from biasing pseudo-label generation while using spatial priors to refine target predictions.

Listen

Deep neural networks have achieved remarkable accuracy in semantic segmentation, which assigns a specific category label to every pixel in an image. However, deploying these models in real-world systems such as autonomous driving remains difficult because models trained in one environment often fail when exposed to new conditions, such as different cities, lighting, or weather. While training models on synthetic, computer-generated data offers an inexpensive alternative to manual image annotation, a large visual domain gap persists between synthetic simulations and real-world environments. The article addresses this challenge by studying unsupervised domain adaptation, which aims to adapt models trained on labeled source data to entirely unlabeled target environments.

The main objective of the article is to demonstrate an effective, self-training framework that aligns synthetic and real data representations without relying on complex adversarial training techniques, while resolving class-imbalance issues in generated pseudo-labels.

To achieve this, the authors develop an iterative self-training approach formulated as a unified loss minimization problem. The algorithm alternately generates high-confidence pseudo-labels for unlabeled target images and updates the neural network using both the original source data and these pseudo-labeled target samples. Because standard self-training naturally favors dominant and easily transferred classes (such as roads or sky) over rarer or harder categories (such as riders, traffic signs, and bikes), the authors introduce Class-Balanced Self-Training (CBST), which normalizes prediction confidence scores independently for each class. Additionally, they incorporate spatial scene priors derived from geometric layout consistencies in driving scenes to further guide label selection. The framework was evaluated across major synthetic-to-real benchmarks (GTA5 to Cityscapes and SYNTHIA to Cityscapes) and cross-city transfers (Cityscapes to NTHU).

The evaluation yielded several key findings. First, self-training combined with modern deep architectures matches or surpasses prevailing adversarial distribution-matching methods without requiring separate discriminator networks. Second, class-balanced selection significantly boosts accuracy on underrepresented and difficult categories; on the SYNTHIA-to-Cityscapes benchmark, CBST improved overall mean Intersection-over-Union (mIoU) to 42.5% (48.4% on a standard 13-class subset), outperforming competing approaches. Third, on the GTA5-to-Cityscapes adaptation task, combining class balancing with spatial priors achieved an mIoU of 46.2%, and 47.0% with multi-scale testing, establishing new state-of-the-art benchmark results. Fourth, the approach demonstrated strong transferability in real-world cross-city testing across Rome, Rio, Tokyo, and Taipei.

These findings indicate that organizations developing vision systems for robotics and autonomous driving can substantially lower the cost and turnaround time of manual data labeling by relying more heavily on synthetic data. By unifying feature alignment and task training into a single optimization process, engineering teams can also simplify model training pipelines and eliminate the training instabilities typical of adversarial architectures.

Decision-makers should consider adopting class-balanced self-training workflows when expanding perception models into new operating domains or geographic markets. For immediate implementation, teams should leverage spatial priors when camera viewpoints and scene geometry remain relatively consistent across environments. Prior to broad deployment, further validation is recommended to evaluate performance under severe environmental variations, such as nighttime conditions or extreme weather, where structural and visual assumptions may diverge significantly.

Cover for Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training

Abstract

Recent deep networks achieved state of the art performance on a variety of semantic segmentation tasks. Despite such progress, these models often face challenges in real world “wild tasks” where large difference between labeled training/source data and unseen test/target data exists. In particular, such difference is often referred to as “domain gap”, and could cause significantly decreased performance which cannot be easily remedied by further increasing the representation power. Unsupervised domain adaptation (UDA) seeks to overcome such problem without target domain labels. In this paper, we propose a novel UDA framework based on an iterative self-training (ST) procedure, where the problem is formulated as latent variable loss minimization, and can be solved by alternatively generating pseudo labels on target data and re-training the model with these labels. On top of ST, we also propose a novel class-balanced self-training (CBST) framework to avoid the gradual dominance of large classes on pseudo-label generation, and introduce spatial priors to refine generated labels. Comprehensive experiments show that the proposed methods achieve state of the art semantic segmentation performance under multiple major UDA settings.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Preliminaries
  • 3.1 Fine-Tuning for Supervised Domain Adaptation
  • 3.2 Self-training for Unsupervised Domain Adaptation
  • 4 Proposed Methods
  • 4.1 Self-training (ST) with Self-paced Learning
  • 4.2 Class-Balanced Self-training (CBST)
  • 4.3 Self-paced Learning Policy Design
  • 4.4 Incorporating Spatial Priors
  • 5 Numerical Experiments
  • 5.1 Small Shift: Cross City Adaptation
  • 5.2 Large Shift: Synthetic to Real Adaptation
  • 6 Conclusions
  • References

Knowls

  1. Knowl 1 — Class-Balanced Self-Training (CBST) Formulation

    model/method

    In unsupervised domain adaptation for semantic segmentation, standard self-training often biases pseudo-label selection toward easy-to-transfer classes that exhibit higher prediction confidences, starving harder classes. Class-Balanced Self-Training (CBST) resolves this issue by assigning class-specific threshold parameters kck_c to normalize confidence scores independently across classes.

    The CBST problem is formulated as a mixed-integer optimization objective:

    min⁡w,y^LCB(w,y^)=−∑s=1S∑n=1Nys,n⊤log⁡(pn(w,Is))−∑t=1T∑n=1N∑c=1C[y^t,n(c)log⁡(pn(c∣w,It))+kcy^t,n(c)]\min_{\mathbf{w}, \hat{\mathbf{y}}} \mathcal{L}_{\mathrm{CB}}(\mathbf{w}, \hat{\mathbf{y}}) = - \sum_{s=1}^S \sum_{n=1}^N \mathbf{y}_{s,n}^\top \log(p_n(\mathbf{w}, \mathbf{I}_s)) - \sum_{t=1}^T \sum_{n=1}^N \sum_{c=1}^C \left[ \hat{y}_{t,n}^{(c)} \log(p_n(c \mid \mathbf{w}, \mathbf{I}_t)) + k_c \hat{y}_{t,n}^{(c)} \right]

    s.t.y^t,n=[y^t,n(1),…,y^t,n(C)]∈{e(i)∣e(i)∈RC}∪{0},∀t,n;kc>0,∀c\text{s.t.} \quad \hat{\mathbf{y}}_{t,n} = [\hat{y}_{t,n}^{(1)}, \dots, \hat{y}_{t,n}^{(C)}] \in \{\mathbf{e}^{(i)} \mid \mathbf{e}^{(i)} \in \mathbb{R}^C\} \cup \{\mathbf{0}\}, \quad \forall t, n; \quad k_c > 0, \quad \forall c

    where Is\mathbf{I}_s is the ss-th source domain image (s∈{1,…,S}s \in \{1, \dots, S\}), It\mathbf{I}_t is the tt-th target domain image (t∈{1,…,T}t \in \{1, \dots, T\}), n∈{1,…,N}n \in \{1, \dots, N\} indexes pixel locations, CC is the number of semantic classes, w\mathbf{w} represents the neural network parameters, ys,n\mathbf{y}_{s,n} is the ground-truth one-hot label vector for pixel nn of source image Is\mathbf{I}_s, and pn(c∣w,I)p_n(c \mid \mathbf{w}, \mathbf{I}) denotes the softmax probability predicted for class cc at pixel nn. The pseudo-label vector y^t,n\hat{\mathbf{y}}_{t,n} is restricted to either a one-hot unit vector e(i)\mathbf{e}^{(i)} (assigning class ii) or a zero vector 0\mathbf{0} (ignoring pixel nn during backpropagation). Each positive scalar kck_c controls the proportion of selected pseudo-labels for class cc.

  2. Knowl 2 — Optimal Pseudo-Label Solver for CBST

    theoretical result

    When network weights w\mathbf{w} and class threshold parameters kck_c are fixed, the pseudo-label optimization step in Class-Balanced Self-Training (CBST) is a discrete optimization problem over all target pixels:

    min⁡y^−∑t=1T∑n=1N∑c=1C[y^t,n(c)log⁡(pn(c∣w,It))+kcy^t,n(c)]\min_{\hat{\mathbf{y}}} - \sum_{t=1}^T \sum_{n=1}^N \sum_{c=1}^C \left[ \hat{y}_{t,n}^{(c)} \log(p_n(c \mid \mathbf{w}, \mathbf{I}_t)) + k_c \hat{y}_{t,n}^{(c)} \right]

    s.t.y^t,n∈{e(i)∣e(i)∈RC}∪{0},∀t,n;kc>0,∀c\text{s.t.} \quad \hat{\mathbf{y}}_{t,n} \in \{\mathbf{e}^{(i)} \mid \mathbf{e}^{(i)} \in \mathbb{R}^C\} \cup \{\mathbf{0}\}, \quad \forall t, n; \quad k_c > 0, \quad \forall c

    The exact closed-form optimal pseudo-label y^t,n(c)∗\hat{y}_{t,n}^{(c)*} for class cc at pixel nn in target image It\mathbf{I}_t is given by:

    y^t,n(c)∗={1,if c=arg⁡max⁡c′pn(c′∣w,It)exp⁡(−kc′)andpn(c∣w,It)exp⁡(−kc)>10,otherwise\hat{y}_{t,n}^{(c)*} = \begin{cases} 1, & \text{if } c = \arg\max_{c'} \frac{p_n(c' \mid \mathbf{w}, \mathbf{I}_t)}{\exp(-k_{c'})} \quad \text{and} \quad \frac{p_n(c \mid \mathbf{w}, \mathbf{I}_t)}{\exp(-k_c)} > 1 \\ 0, & \text{otherwise} \end{cases}

    Under this decision rule, pseudo-label assignment depends on the normalized probability ratio pn(c∣w,It)/exp⁡(−kc)p_n(c \mid \mathbf{w}, \mathbf{I}_t) / \exp(-k_c). A pixel is assigned a pseudo-label only when its normalized response exceeds 11; if multiple classes satisfy this condition, the class with the maximum normalized score is selected.

  3. Knowl 3 — Class-Wise Threshold Determination Algorithm for CBST

    algorithm

    To dynamically adjust the class-specific parameters kck_c in Class-Balanced Self-Training (CBST), a self-paced curriculum is employed. For each class cc, the algorithm identifies all target domain pixels predicted as class cc, sorts their predicted probabilities in descending order, and chooses kck_c such that exp⁡(−kc)\exp(-k_c) matches the probability at rank ⌊p×Nc⌋\lfloor p \times N_c \rfloor, where NcN_c is the total number of pixels predicted as class cc across all target images.

    Input : Neural network f(w), all target images It (t = 1 to T), selection proportion p
    Output: Vector of class threshold parameters kc for c = 1 to C
    for t = 1 to T do
        PIt = P(w, It)
        LPIt = argmax(PIt, axis=0)
        MPIt = max(PIt, axis=0)
        for c = 1 to C do
            MPc,It = MPIt[LPIt == c]
            Mc = [Mc, matrix_to_vector(MPc,It)]
        end
    end
    for c = 1 to C do
        Mc = sort(Mc, order=descending)
        lenc,th = length(Mc) * p
        kc = -log(Mc[lenc,th])
    end
    return kc

    The proportion parameter pp controls the fraction of pseudo-labeled pixels per class. In training, pp is initialized at 20%20\%, increased by 5%5\% at each subsequent self-training round, and capped at a maximum of 50%50\%.

  4. Knowl 4 — Class-Balanced Self-Training with Spatial Priors (CBST-SP)

    model/method

    For urban street scene adaptation where camera viewpoints and spatial layout are consistent across domains, spatial priors qn(c)q_n(c) provide structural regularization for pseudo-label generation. The spatial prior qn(c)q_n(c) denotes the normalized spatial frequency of class cc at pixel location nn across all source domain training images, smoothed with a 70×7070 \times 70 Gaussian kernel and normalized such that ∑n=1Nqn(c)=1\sum_{n=1}^N q_n(c) = 1.

    The spatial prior is integrated into the Class-Balanced Self-Training objective by weighting the target class probability pn(c∣w,It)p_n(c \mid \mathbf{w}, \mathbf{I}_t) by qn(c)q_n(c) during pseudo-label generation:

    min⁡w,y^LSP(w,y^)=−∑s=1S∑n=1Nys,n⊤log⁡(pn(w,Is))−∑t=1T∑n=1N∑c=1C[y^t,n(c)log⁡(qn(c)pn(c∣w,It))+kcy^t,n(c)]\min_{\mathbf{w}, \hat{\mathbf{y}}} \mathcal{L}_{\mathrm{SP}}(\mathbf{w}, \hat{\mathbf{y}}) = - \sum_{s=1}^S \sum_{n=1}^N \mathbf{y}_{s,n}^\top \log(p_n(\mathbf{w}, \mathbf{I}_s)) - \sum_{t=1}^T \sum_{n=1}^N \sum_{c=1}^C \left[ \hat{y}_{t,n}^{(c)} \log(q_n(c) p_n(c \mid \mathbf{w}, \mathbf{I}_t)) + k_c \hat{y}_{t,n}^{(c)} \right]

    s.t.y^t,n∈{e(i)∣e(i)∈RC}∪{0},∀t,n;kc>0,∀c\text{s.t.} \quad \hat{\mathbf{y}}_{t,n} \in \{\mathbf{e}^{(i)} \mid \mathbf{e}^{(i)} \in \mathbb{R}^C\} \cup \{\mathbf{0}\}, \quad \forall t, n; \quad k_c > 0, \quad \forall c

    Because qn(c)q_n(c) is independent of network weights w\mathbf{w}, it acts as an additive constant inside log⁡(qn(c)pn(c∣w,It))=log⁡qn(c)+log⁡pn(c∣w,It)\log(q_n(c) p_n(c \mid \mathbf{w}, \mathbf{I}_t)) = \log q_n(c) + \log p_n(c \mid \mathbf{w}, \mathbf{I}_t) during network backpropagation; thus, spatial priors modulate the selection of target pseudo-labels y^\hat{\mathbf{y}} without altering the gradient update formula for w\mathbf{w}.

  5. Knowl 5 — Self-Training (ST) with Self-Paced Learning Formulation

    model/method

    Unsupervised domain adaptation for semantic segmentation via basic self-training frames target pseudo-label generation and model fine-tuning as an alternating block coordinate descent minimization over network weights w\mathbf{w} and discrete target pseudo-labels y^\hat{\mathbf{y}}. An L1L_1 penalty on y^\hat{\mathbf{y}} acts as a negative sparsity-promoting term to prevent the trivial solution of discarding all target pseudo-labels:

    min⁡w,y^LST(w,y^)=−∑s=1S∑n=1Nys,n⊤log⁡(pn(w,Is))−∑t=1T∑n=1N[y^t,n⊤log⁡(pn(w,It))+k∥y^t,n∥1]\min_{\mathbf{w}, \hat{\mathbf{y}}} \mathcal{L}_{\mathrm{ST}}(\mathbf{w}, \hat{\mathbf{y}}) = - \sum_{s=1}^S \sum_{n=1}^N \mathbf{y}_{s,n}^\top \log(p_n(\mathbf{w}, \mathbf{I}_s)) - \sum_{t=1}^T \sum_{n=1}^N \left[ \hat{\mathbf{y}}_{t,n}^\top \log(p_n(\mathbf{w}, \mathbf{I}_t)) + k \|\hat{\mathbf{y}}_{t,n}\|_1 \right]

    s.t.y^t,n∈{e(i)∣e(i)∈RC}∪{0},∀t,n;k>0\text{s.t.} \quad \hat{\mathbf{y}}_{t,n} \in \{\mathbf{e}^{(i)} \mid \mathbf{e}^{(i)} \in \mathbb{R}^C\} \cup \{\mathbf{0}\}, \quad \forall t, n; \quad k > 0

    Alternating optimization consists of two recurring steps per round:

    1. Fixing w\mathbf{w}, the pseudo-label for each target pixel is assigned via the closed-form rule:

    y^t,n(c)∗={1,if c=arg⁡max⁡c′pn(c′∣w,It)andpn(c∣w,It)>exp⁡(−k)0,otherwise\hat{y}_{t,n}^{(c)*} = \begin{cases} 1, & \text{if } c = \arg\max_{c'} p_n(c' \mid \mathbf{w}, \mathbf{I}_t) \quad \text{and} \quad p_n(c \mid \mathbf{w}, \mathbf{I}_t) > \exp(-k) \\ 0, & \text{otherwise} \end{cases}

    1. Fixing y^\hat{\mathbf{y}}, the network parameters w\mathbf{w} are updated by minimizing standard cross-entropy loss over both labeled source data and pseudo-labeled target data using stochastic gradient descent.
  6. Knowl 6 — Global Threshold Determination Algorithm for Standard Self-Training

    algorithm

    In standard self-training (ST) with self-paced learning, a single scalar parameter kk determines the cutoff above which a target pixel's maximum class probability is selected as a valid pseudo-label. The threshold is selected so that a fixed top proportion pp of all pixels across the entire target dataset are retained.

    Input : Neural network P(w), all target images It (t = 1 to T), selection proportion p
    Output: Global threshold parameter k
    for t = 1 to T do
        PIt = P(w, It)
        MPIt = max(PIt, axis=0)
        M = [M, matrix_to_vector(MPIt)]
    end
    M = sort(M, order=descending)
    lenth = length(M) * p
    k = -log(M[lenth])
    return k

    The proportion pp follows a self-paced schedule: starting at 20%20\%, adding 5%5\% at each self-training round, and stopping at a ceiling of 50%50\%.

  7. Knowl 7 — Synthetic-to-Real Domain Adaptation Benchmark on GTA5 to Cityscapes

    data/table

    Adapting segmentation models from synthetic GTA5 images (24,96624,966 annotated images of size 1052×19141052 \times 1914) to real Cityscapes validation images (500500 images) across 1919 evaluation categories demonstrates that CBST and CBST with spatial priors (CBST-SP) achieve superior adaptation performance over standard self-training (ST) and adversarial adaptation baselines.

    Method Base Net Road SW Build Wall Fence Pole TL TS Veg. Terrain Sky PR Rider Car Truck Bus Train Motor Bike mIoU
    Source only FCN8s-VGG16 64.0 22.1 68.6 13.3 8.7 19.9 15.5 5.9 74.9 13.4 37.0 37.7 10.3 48.2 6.1 1.2 1.8 10.8 2.9 24.3
    ST FCN8s-VGG16 83.8 17.4 72.1 14.3 2.9 16.5 16.0 6.8 81.4 24.2 47.2 40.7 7.6 71.7 10.2 7.6 0.5 11.1 0.9 28.1
    CBST FCN8s-VGG16 66.7 26.8 73.7 14.8 9.5 28.3 25.9 10.1 75.5 15.7 51.6 47.2 6.2 71.9 3.7 2.2 5.4 18.9 32.4 30.9
    CBST-SP FCN8s-VGG16 90.4 50.8 72.0 18.3 9.5 27.2 28.6 14.1 82.4 25.1 70.8 42.6 14.5 76.9 5.9 12.5 1.2 14.0 28.6 36.1
    CyCADA FCN8s-VGG16 85.2 37.2 76.5 21.8 15.0 23.8 22.9 21.5 80.5 31.3 60.7 50.5 9.0 76.9 17.1 28.2 4.5 9.8 0.0 35.4
    Source only ResNet-38 70.0 23.7 67.8 15.4 18.1 40.2 41.9 25.3 78.8 11.7 31.4 62.9 29.8 60.1 21.5 26.8 7.7 28.1 12.0 35.4
    ST ResNet-38 90.1 56.8 77.9 28.5 23.0 41.5 45.2 39.6 84.8 26.4 49.2 59.0 27.4 82.3 39.7 45.6 20.9 34.8 46.2 41.5
    CBST ResNet-38 86.8 46.7 76.9 26.3 24.8 42.0 46.0 38.6 80.7 15.7 48.0 57.3 27.9 78.2 24.5 49.6 17.7 25.5 45.1 45.2
    CBST-SP ResNet-38 88.0 56.2 77.0 27.4 22.4 40.7 47.3 40.9 82.4 21.6 60.3 50.2 20.4 83.8 35.0 51.0 15.2 20.6 37.0 46.2
    CBST-SP+MST ResNet-38 89.6 58.9 78.5 33.0 22.3 41.4 48.2 39.2 83.6 24.3 65.4 49.3 20.2 83.3 39.0 48.6 12.5 20.3 35.3 47.0

    On FCN8s-VGG16, CBST-SP achieves 36.1%36.1\% mIoU, outperforming CyCADA (35.4%35.4\%). On ResNet-38, CBST-SP reaches 46.2%46.2\% mIoU, and with multi-scale testing (MST at scales 0.5,0.75,1.00.5, 0.75, 1.0) reaches 47.0%47.0\% mIoU.

  8. Knowl 8 — Synthetic-to-Real Domain Adaptation Benchmark on SYNTHIA to Cityscapes

    data/table

    Models trained on the SYNTHIA-RAND-CITYSCAPES subset (9,4009,400 images of size 760×1280760 \times 1280) were adapted to Cityscapes validation images. Evaluation reports mean IoU over 16 common classes (mIoU) and over 13 common classes (mIoU*, excluding wall, fence, and pole).

    Method Base Net Road SW Build Wall* Fence* Pole* TL TS Veg. Sky PR Rider Car Bus Motor Bike mIoU mIoU*
    Source only FCN8s-VGG16 17.2 19.7 47.3 1.1 0.0 19.1 3.0 9.1 71.8 78.3 37.6 4.7 42.2 9.0 0.1 0.9 22.6 26.2
    ST FCN8s-VGG16 0.2 14.5 53.8 1.6 0.0 18.9 0.9 7.8 72.2 80.3 48.1 6.3 67.7 4.7 0.2 4.5 23.9 27.8
    CBST FCN8s-VGG16 69.6 28.7 69.5 12.1 0.1 25.4 11.9 13.6 82.0 81.9 49.1 14.5 66.0 6.6 3.7 32.4 35.4 36.1
    Curr. DA FCN8s-VGG16 65.2 26.1 74.9 0.1 0.5 10.7 3.5 3.0 76.1 70.6 47.1 8.2 43.2 20.7 0.7 13.1 29.0 34.8
    GAN DA FCN8s-VGG16 79.1 31.1 77.1 3.0 0.2 22.8 6.6 15.2 77.4 78.9 47.0 14.8 67.5 16.3 6.9 13.0 34.8 40.8
    Source only ResNet-38 32.6 21.5 46.5 4.8 0.1 26.5 14.8 13.1 70.8 60.3 56.6 3.5 74.1 20.4 8.9 13.1 29.2 33.6
    ST ResNet-38 38.2 19.6 70.2 3.9 0.0 31.9 17.6 17.2 82.4 68.3 63.1 5.3 78.4 11.2 0.8 7.5 32.2 36.9
    CBST ResNet-38 53.6 23.7 75.0 12.5 0.3 36.4 23.5 26.3 84.8 74.7 67.2 17.5 84.5 28.4 15.2 55.8 42.5 48.4

    CBST prevents standard ST from degenerating on hard-to-transfer classes (e.g., road, wall, rider, motorcycle, bicycle). With ResNet-38, CBST achieves 42.5%42.5\% 16-class mIoU and 48.4%48.4\% 13-class mIoU*, improving over MAA (46.7%46.7\% on 13 classes).

  9. Knowl 9 — Cross-City Domain Adaptation Benchmark on Cityscapes to NTHU

    data/table

    Cross-city adaptation evaluates transfer from Cityscapes training data to the NTHU dataset (400400 images of resolution 1024×20481024 \times 2048 spanning four target cities: Rome, Rio, Tokyo, and Taipei) across 13 shared classes using a 10-fold cross-validation protocol.

    City Method Road SW Build TL TS Veg. Sky PR Rider Car Bus Motor Bike Mean
    Rome Source ResNet-38 86.0 21.4 81.5 14.3 47.4 82.9 59.8 30.8 20.9 83.1 20.2 40.0 5.6 45.7
    ST 85.9 20.2 84.3 15.0 46.4 84.9 73.5 48.5 21.6 84.6 17.6 46.2 6.7 48.9
    CBST 87.1 43.9 89.7 14.8 47.7 85.4 90.3 45.4 26.6 85.4 20.5 49.8 10.3 53.6
    Rio Source ResNet-38 80.6 36.0 81.8 21.0 33.1 79.0 64.7 36.0 21.0 73.1 33.6 22.5 7.8 45.4
    ST 80.1 41.4 83.8 19.1 39.1 80.8 71.2 56.3 27.7 79.9 32.7 36.4 12.2 50.8
    CBST 84.3 55.2 85.4 19.6 30.1 80.5 77.9 55.2 28.6 79.7 33.2 37.6 11.5 52.2
    Tokyo Source ResNet-38 83.8 26.4 73.0 6.5 27.0 80.5 46.6 35.6 22.8 71.3 4.2 10.5 36.1 40.3
    ST 83.1 27.7 74.8 7.1 29.4 84.4 48.5 57.2 23.3 73.3 3.3 22.7 45.8 44.6
    CBST 85.2 33.6 80.4 8.3 31.1 83.9 78.2 53.2 28.9 72.7 4.4 27.0 47.0 48.8
    Taipei Source ResNet-38 84.9 26.0 80.1 8.3 28.0 73.9 54.4 18.9 26.8 71.6 26.0 48.2 14.7 43.2
    ST 83.1 23.5 78.2 9.6 25.4 74.8 35.9 33.2 27.3 75.2 32.3 52.2 28.8 44.6
    CBST 86.1 35.2 84.2 15.0 22.2 75.6 74.9 22.7 33.1 78.0 37.6 58.0 30.9 50.3

    CBST with ResNet-38 achieves 53.6%53.6\% (Rome), 52.2%52.2\% (Rio), 48.8%48.8\% (Tokyo), and 50.3%50.3\% (Taipei) mIoU, consistently outperforming source-only models and standard ST across all four target cities.

  10. Knowl 10 — Hard Sample Mining Strategy for Underrepresented Target Classes

    model/method

    In Class-Balanced Self-Training (CBST) and CBST with Spatial Priors (CBST-SP) on challenging adaptation benchmarks (GTA5 →\to Cityscapes and Cityscapes →\to NTHU), a hard sample mining heuristic is applied during pseudo-label selection to prevent rare classes from vanishing. The strategy identifies the 55 classes with the lowest predicted pixel frequencies in the target domain, granting top selection priority to any class whose total predicted portion occupies less than 0.1%0.1\% of all target pixels.

Coverage note — None was omitted; all key algorithmic contributions (ST, CBST, CBST-SP, threshold determination algorithms, hard class mining) and primary benchmark results on GTA5, SYNTHIA, and NTHU datasets are fully captured.

References

  1. 1.http://data-bdd.berkeley.edu/
  2. 2.Bekker, A.J., Goldberger, J.: Training deep neural-networks based on unreliable labels. In: 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2682–2686. IEEE (2016)
  3. 3.Bousmalis, K., Trigeorgis, G., Silberman, N., Krishnan, D., Erhan, D.: Domain separation networks. In: Advances in Neural Information Processing Systems, pp. 343–351 (2016)
  4. 4.Chapelle, O., Scholkopf, B., Zien, A.: Semi-supervised learning. IEEE Trans. Neural Netw. 20(3), 542–542 (2009). (Chapelle, O. et al. (eds.) 2006) [book reviews]
  5. 5.Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: DeepLab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. arXiv preprint arXiv:1606.00915 (2016)
  6. 6.Chen, L.C., Papandreou, G., Schroff, F., Adam, H.: Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587 (2017)
  7. 7.Chen, M., Weinberger, K.Q., Blitzer, J.: Co-training for domain adaptation. In: Advances in Neural Information Processing Systems, pp. 2456–2464 (2011)
  8. 8.Chen, T., et al.: MXNet: a flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274 (2015)
  9. 9.Chen, Y.H., Chen, W.Y., Chen, Y.T., Tsai, B.C., Frank Wang, Y.C., Sun, M.: No more discrimination: cross city adaptation of road scene segmenters. In: The IEEE International Conference on Computer Vision (ICCV), October 2017
  10. 10.Cordts, M., et al.: The cityscapes dataset for semantic urban scene understanding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3213–3223 (2016)
  11. 11.Ganin, Y., Lempitsky, V.: Unsupervised domain adaptation by backpropagation. In: International Conference on Machine Learning, pp. 1180–1189 (2015)
  12. 12.Ganin, Y., et al.: Domain-adversarial training of neural networks. J. Mach. Learn. Res. 17(59), 1–35 (2016)
  13. 13.Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving? The KITTI vision benchmark suite. In: 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3354–3361. IEEE (2012)
  14. 14.Gretton, A., Smola, A., Huang, J., Schmittfull, M., Borgwardt, K., Sch¨olkopf, B.: Covariate shift and local learning by distribution matching, pp. 131–160. MIT Press, Cambridge (2009)
  15. 15.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778 (2016)
  16. 16.Hoffman, J., et al.: CyCADA: cycle-consistent adversarial domain adaptation. arXiv preprint arXiv:1711.03213 (2017)
  17. 17.Hoffman, J., Wang, D., Yu, F., Darrell, T.: FCNs in the wild: pixel-level adversarial and constraint-based adaptation. arXiv preprint arXiv:1612.02649 (2016)
  18. 18.Huang, G., Liu, Z.: Densely connected convolutional networks
  19. 19.Krizhevsky, A., Sutskever, I., Hinton, G.E.: ImageNet classification with deep convolutional neural networks. In: Advances in Neural Information Processing Systems (2012)
  20. 20.Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3431–3440 (2015)
  21. 21.Long, M., Cao, Y., Wang, J., Jordan, M.: Learning transferable features with deep adaptation networks. In: International Conference on Machine Learning, pp. 97–105 (2015)
  22. 22.Maeireizo, B., Litman, D., Hwa, R.: Co-training for predicting emotions with spoken dialogue data. In: Proceedings of the ACL 2004 on Interactive Poster and Demonstration Sessions, p. 28. Association for Computational Linguistics (2004)
  23. 23.Murez, Z., Kolouri, S., Kriegman, D., Ramamoorthi, R., Kim, K.: Image to image translation for domain adaptation. arXiv preprint arXiv:1712.00479 (2017)
  24. 24.Richter, S.R., Vineet, V., Roth, S., Koltun, V.: Playing for data: ground truth from computer games. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9906, pp. 102–118. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46475-6 7
  25. 25.Riloff, E., Wiebe, J., Wilson, T.: Learning subjective nouns using extraction pattern bootstrapping. In: Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, vol. 4, pp. 25–32. Association for Computational Linguistics (2003)
  26. 26.Ros, G., Sellart, L., Materzynska, J., Vazquez, D., Lopez, A.M.: The SYNTHIA dataset: a large collection of synthetic images for semantic segmentation of urban scenes. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3234–3243 (2016)
  27. 27.Russakovsky, O., et al.: Imagenet large scale visual recognition challenge. Int. J. Comput. Vis. 115(3), 211–252 (2015)
  28. 28.Saito, K., Ushiku, Y., Harada, T., Saenko, K.: Adversarial dropout regularization. arXiv preprint arXiv:1711.01575 (2017)
  29. 29.Sankaranarayanan, S., Balaji, Y., Jain, A., Lim, S.N., Chellappa, R.: Unsupervised domain adaptation for semantic segmentation with gans. arXiv preprint arXiv:1711.06969 (2017)
  30. 30.Silberman, N., Fergus, R.: Indoor scene segmentation using a structured light sensor. In: 2011 IEEE International Conference on Computer Vision Workshops (ICCV Workshops), pp. 601–608. IEEE (2011)
  31. 31.Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. In: International Conference on Learning Representations (2015)
  32. 32.Sun, B., Saenko, K.: Deep CORAL: correlation alignment for deep domain adaptation. In: Hua, G., J´egou, H. (eds.) ECCV 2016. LNCS, vol. 9915, pp. 443–450. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-49409-8 35
  33. 33.Tang, K., Ramanathan, V., Fei-Fei, L., Koller, D.: Shifting weights: adapting object detectors from image to video. In: Advances in Neural Information Processing Systems, pp. 638–646 (2012)
  34. 34.Tsai, Y.H., Hung, W.C., Schulter, S., Sohn, K., Yang, M.H., Chandraker, M.: Learning to adapt structured output space for semantic segmentation. arXiv preprint arXiv:1802.10349 (2018)
  35. 35.Tzeng, E., Hoffman, J., Darrell, T., Saenko, K.: Simultaneous deep transfer across domains and tasks. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 4068–4076 (2015)
  36. 36.Tzeng, E., Hoffman, J., Darrell, T., Saenko, K.: Adversarial discriminative domain adaptation. In: Computer Vision and Pattern Recognition (CVPR) (2017)
  37. 37.Tzeng, E., Hoffman, J., Zhang, N., Saenko, K., Darrell, T.: Deep domain confusion: maximizing for domain invariance. arXiv preprint arXiv:1412.3474 (2014)
  38. 38.Wang, P., et al.: Understanding convolution for semantic segmentation. arXiv preprint arXiv:1702.08502 (2017)
  39. 39.Wu, Z., Shen, C., van den Hengel, A.: Wider or deeper: revisiting the ResNet model for visual recognition. arXiv preprint arXiv:1611.10080 (2016)
  40. 40.Yarowsky, D.: Unsupervised word sense disambiguation rivaling supervised methods. In: Proceedings of the 33rd Annual Meeting on Association for Computational Linguistics, pp. 189–196. Association for Computational Linguistics (1995)
  41. 41.Yu, F., Koltun, V.: Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122 (2015)
  42. 42.Yu, F., Koltun, V., Funkhouser, T.: Dilated residual networks. In: Computer Vision and Pattern Recognition, vol. 1 (2017)
  43. 43.Zhang, Y., David, P., Gong, B.: Curriculum domain adaptation for semantic segmentation of urban scenes. In: The IEEE International Conference on Computer Vision (ICCV), October 2017
  44. 44.Zhao, H., Shi, J., Qi, X., Wang, X., Jia, J.: Pyramid scene parsing network. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017
  45. 45.Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., Torralba, A.: Semantic understanding of scenes through the ADE20K dataset. arXiv preprint arXiv:1608.05442 (2016)
  46. 46.Zhu, J.Y., Park, T., Isola, P., Efros, A.A.: Unpaired image-to-image translation using cycle-consistent adversarial networks
  47. 47.Zhu, X.: Semi-supervised learning literature survey (2005)

Citation

MLA
Zou, Y., et al. “Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training”. Lecture Notes in Computer Science, Springer International Publishing, 2018, pp. 297–313, https://doi.org/10.1007/978-3-030-01219-9_18.
APA
Zou, Y., Yu, Z., Vijaya Kumar, B. V. K., & Wang, J. (2018). Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training. In Lecture Notes in Computer Science (pp. 297–313). Springer International Publishing. https://doi.org/10.1007/978-3-030-01219-9_18
Chicago
Zou, Y., Z. Yu, B. V. K. Vijaya Kumar, and J. Wang. 2018. “Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training”. In Lecture Notes in Computer Science. Springer International Publishing. https://doi.org/10.1007/978-3-030-01219-9_18.
Harvard
Zou, Y. et al. (2018) “Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training”, Lecture Notes in Computer Science. Springer International Publishing, pp. 297–313. Available at: https://doi.org/10.1007/978-3-030-01219-9_18.
Vancouver
1. Zou Y, Yu Z, Vijaya Kumar BVK, Wang J (2018) Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training. In: Lecture Notes in Computer Science. Springer International Publishing, pp 297–313

BibTeX

@inbook{Zou_2018, title={Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training}, ISBN={9783030012199}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/978-3-030-01219-9_18}, DOI={10.1007/978-3-030-01219-9_18}, booktitle={Computer Vision – ECCV 2018}, publisher={Springer International Publishing}, author={Zou, Yang and Yu, Zhiding and Vijaya Kumar, B. V. K. and Wang, Jinsong}, year={2018}, pages={297–313} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF