Contrastive Test-Time Adaptation

Dian ChenDequan WangTrevor DarrellSayna Ebrahimi

article2022CVPR373 citations

Develops AdaContrast, a test-time adaptation framework that pairs contrastive representation learning with nearest-neighbor pseudo-label refinement to achieve state-of-the-art target accuracy and well-calibrated predictions without requiring source data.

Listen

Machine learning vision models often suffer severe performance drops when deployed to new operating environments that differ from their training data. Addressing this domain shift typically requires access to original training data, which raises significant data privacy and bandwidth concerns. The article develops and evaluates AdaContrast, a test-time adaptation method that enables visual classification models to adapt directly to unlabeled target data without accessing original source data.

The approach combines joint self-supervised contrastive learning with an online pseudo-label refinement mechanism. Rather than relying on computationally heavy generative models or error-prone offline label updates, AdaContrast refines predictions per batch using soft voting among nearest neighbors in a target feature memory queue. In parallel, it initializes target feature encoders using source model weights and trains a contrastive objective that discards same-class negative pairs, while applying consistency and class diversification regularizations. Experiments were conducted on standard image classification benchmarks, including VisDA-C and the 126-class DomainNet dataset across seven domain shifts.

The evaluation demonstrates that AdaContrast establishes new state-of-the-art performance in source-free domain adaptation. On the VisDA-C benchmark, the method achieved an 86.8% average accuracy, outperforming the previous best source-free approach by 3.8 percentage points. On DomainNet-126, it delivered an average accuracy of 67.8% across seven shifts, outperforming competing test-time approaches by up to 10.1 percentage points and even surpassing standard adaptation methods that require source data. Furthermore, AdaContrast achieved substantial gains in model calibration, reducing expected calibration error by a factor of 4.5 compared to prior entropy minimization techniques. The method also proved stable across varying learning rates and retained high accuracy using a compact memory queue containing under 4% of target dataset features.

These findings indicate that organizations can safely and cost-effectively update artificial intelligence models at the edge or in streaming environments without transferring proprietary or private source datasets. By avoiding overconfident, miscalibrated predictions, the framework reduces operational risk in downstream decision systems. For deployment, engineering teams should consider adopting AdaContrast for streaming applications, such as robotics, while taking care to monitor model trustworthiness and potential misuse in safety-critical pipelines.

No sufficiently relevant recommendations were found.

Cover for Contrastive Test-Time Adaptation

Abstract

Test-time adaptation is a special setting of unsupervised domain adaptation where a trained model on the source domain has to adapt to the target domain without accessing source data. We propose a novel way to leverage self-supervised contrastive learning to facilitate target feature learning, along with an online pseudo labeling scheme with refinement that significantly denoises pseudo labels. The contrastive learning task is applied jointly with pseudo labeling, contrasting positive and negative pairs constructed similarly as MoCo but with source-initialized encoder, and excluding same-class negative pairs indicated by pseudo labels. Meanwhile, we produce pseudo labels online and refine them via soft voting among their nearest neighbors in the target feature space, enabled by maintaining a memory queue. Our method, AdaContrast, achieves state-of-the-art performance on major benchmarks while having several desirable properties compared to existing works, including memory efficiency, insensitivity to hyper-parameters, and better model calibration. Project page: this http URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Online pseudo label refinement
  • 3.2 Joint self-supervised contrastive learning
  • 3.3 Additional regularization
  • 4 Experiments
  • 4.1 Experimental setup
  • 4.2 Results
  • 4.3 Analysis and Discussion
  • 4.4 Ablation studies
  • 5 Limitations
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — AdaContrast Framework for Test-Time Adaptation

    model/method

    AdaContrast is a test-time adaptation (TTA) framework for closed-set image classification that adapts a source-trained model to an unlabeled target domain Xt\mathcal{X}_t without accessing source data Ds\mathcal{D}_s. The target model gt(⋅)=ht(ft(⋅))g_t(\cdot) = h_t(f_t(\cdot)) (comprising a feature extractor ft:Xt→RDf_t: \mathcal{X}_t \to \mathbb{R}^D and a classifier ht:RD→RCh_t: \mathbb{R}^D \to \mathbb{R}^C for CC classes) and a slowly-updating momentum model gt′(⋅)=ht′(ft′(⋅))g'_t(\cdot) = h'_t(f'_t(\cdot)) are initialized with the trained source model parameters θs\theta_s.

    For each unlabeled target input xt∈Xtx_t \in \mathcal{X}_t, AdaContrast constructs three views: one weakly-augmented view tw(xt)t_w(x_t) and two strongly-augmented views ts(xt),ts′(xt)t_s(x_t), t'_s(x_t). The weakly-augmented sample generates an online refined pseudo label y^\hat{y} via soft nearest-neighbor voting against a feature memory queue. The model is optimized jointly using a multi-task objective combining weak-strong self-training cross-entropy Ltce\mathcal{L}^{ce}_t, target self-supervised contrastive learning Ltctr\mathcal{L}^{ctr}_t, and class diversity regularization Ltdiv\mathcal{L}^{div}_t:

    Lt=γ1Ltce+γ2Ltctr+γ3Ltdiv\mathcal{L}_t = \gamma_1 \mathcal{L}^{ce}_t + \gamma_2 \mathcal{L}^{ctr}_t + \gamma_3 \mathcal{L}^{div}_t

    where γ1=γ2=γ3=1.0\gamma_1 = \gamma_2 = \gamma_3 = 1.0 is the default setting.

  2. Knowl 2 — Online Pseudo-Label Refinement via Nearest-Neighbor Soft Voting

    model/method

    To prevent noisy pseudo labels from accumulating errors during test-time adaptation, AdaContrast refines pseudo labels on a per-batch basis via soft kk-nearest neighbor voting in target feature space instead of using offline epoch-level updates.

    A memory queue QwQ_w of length MM stores feature vectors and predicted class probability distributions {w′j,p′j}j=1M\{w'^j, p'^j\}_{j=1}^M of weakly augmented target samples. Features and probabilities in QwQ_w are generated by a momentum model gt′(⋅)=ht′(ft′(⋅))g'_t(\cdot) = h'_t(f'_t(\cdot)):

    w′=ft′(tw(xt)),p′=σ(ht′(w′))w' = f'_t(t_w(x_t)), \quad p' = \sigma(h'_t(w'))

    where σ(⋅)\sigma(\cdot) denotes the softmax function. The momentum model weights θt′\theta'_t are initialized to source model weights θs\theta_s and updated at each mini-batch step with momentum coefficient mm:

    θt′←mθt′+(1−m)θt\theta'_t \leftarrow m\theta'_t + (1-m)\theta_t

    Given a target sample xtx_t and its weakly augmented feature w=ft(tw(xt))w = f_t(t_w(x_t)), the NN nearest neighbors IiI_i are retrieved from QwQ_w using cosine similarity. The refined class probability p^(i,c)\hat{p}^{(i,c)} for class cc is computed by soft voting:

    p^(i,c)=1N∑j∈Iip′(j,c)\hat{p}^{(i,c)} = \frac{1}{N} \sum_{j \in I_i} p'^{(j,c)}

    The refined pseudo label y^i\hat{y}^i assigned to xtx_t is determined by:

    y^i=arg⁡max⁡cp^(i,c)\hat{y}^i = \arg\max_c \hat{p}^{(i,c)}

  3. Knowl 3 — Target Contrastive Loss with Same-Class Negative Pair Exclusion

    equation

    In AdaContrast, self-supervised contrastive learning operates on the target domain using two strongly-augmented views ts(xt)t_s(x_t) and ts′(xt)t'_s(x_t) for each input image xtx_t. The query feature is computed by the student encoder q=ft(ts(xt))q = f_t(t_s(x_t)), and the positive key feature is computed by the momentum encoder k+=ft′(ts′(xt))k^+ = f'_t(t'_s(x_t)). A memory queue QsQ_s of length PP stores past key features and their corresponding refined pseudo labels {kj,y^j}j=1P\{k^j, \hat{y}^j\}_{j=1}^P.

    To prevent pushing apart samples that belong to the same underlying semantic class, negative pairs in QsQ_s that share the same pseudo label as the query (y=y^jy = \hat{y}^j) are excluded from the contrastive denominator. The modified InfoNCE loss is defined as:

    Ltctr=−log⁡exp⁡(q⋅k+/τ)∑j∈Nqexp⁡(q⋅kj/τ)\mathcal{L}^{ctr}_t = -\log \frac{\exp(q \cdot k^+ / \tau)}{\sum_{j \in \mathcal{N}_q} \exp(q \cdot k_j / \tau)}

    where τ\tau is the temperature hyperparameter, k0=k+k_0 = k^+, and the active index set Nq\mathcal{N}_q is:

    Nq={j∣1≤j≤P, j∈Z, y^≠y^j}∪{0}\mathcal{N}_q = \{j \mid 1 \le j \le P,\, j \in \mathbb{Z},\, \hat{y} \ne \hat{y}^j\} \cup \{0\}

  4. Knowl 4 — Weak-Strong Consistency and Diversity Regularization Objectives

    equation

    AdaContrast incorporates weak-strong consistency and class diversity regularization to train the target model without requiring confidence thresholding.

    The self-training consistency loss applies cross-entropy to the strongly-augmented query image predictions using the refined pseudo label y^\hat{y} obtained from the weakly-augmented image view:

    Ltce=−Ext∈Xt∑c=1Cy^clog⁡pqc\mathcal{L}^{ce}_t = -\mathbb{E}_{x_t \in \mathcal{X}_t} \sum_{c=1}^C \hat{y}^c \log p_q^c

    where y^c\hat{y}^c is the one-hot representation of y^\hat{y}, and pq=σ(gt(ts(xt)))p_q = \sigma(g_t(t_s(x_t))) represents the softmax probability output of the target model for the strongly augmented view ts(xt)t_s(x_t).

    To prevent the classifier from collapsing into predicting only a subset of dominant classes, a diversity regularization loss is applied over the batch-averaged prediction distribution:

    Ltdiv=Ext∈Xt∑c=1Cpˉqclog⁡pˉqc\mathcal{L}^{div}_t = \mathbb{E}_{x_t \in \mathcal{X}_t} \sum_{c=1}^C \bar{p}_q^c \log \bar{p}_q^c

    pˉq=Ext∈Xtσ(gt(ts(xt)))\bar{p}_q = \mathbb{E}_{x_t \in \mathcal{X}_t} \sigma(g_t(t_s(x_t)))

  5. Knowl 5 — AdaContrast Adaptation Algorithm

    algorithm

    The overall adaptation loop of AdaContrast optimizes the target encoder and classifier on unlabeled batches using online refined pseudo labels, same-class-filtered momentum contrast, and diversity regularization.

    Input: Target dataset Xt\mathcal{X}_t, source model parameters θs\theta_s, queue lengths MM and PP, momentum factor mm, neighbor count NN, temperature τ\tau.
    Output: Adapted target model gtg_t.
    Initialize target model gt(⋅)=ht(ft(⋅))g_t(\cdot) = h_t(f_t(\cdot)) with parameters θt←θs\theta_t \leftarrow \theta_s
    Initialize momentum model gt′(⋅)=ht′(ft′(⋅))g'_t(\cdot) = h'_t(f'_t(\cdot)) with parameters θt′←θs\theta'_t \leftarrow \theta_s
    Initialize queue QwQ_w of size MM with {w′j,p′j}j=1M\{w'^j, p'^j\}_{j=1}^M from random target samples
    Initialize queue QsQ_s of size PP with {kj,y^j}j=1P\{k^j, \hat{y}^j\}_{j=1}^P from random target samples
    for each mini-batch B⊂XtB \subset \mathcal{X}_t do
        for each xt∈Bx_t \in B do
            Sample augmentations: one weak twt_w, two strong ts,ts′t_s, t'_s
            Compute weak feature w=ft(tw(xt))w = f_t(t_w(x_t))
            Find NN nearest neighbors II of ww in QwQ_w by cosine similarity
            Compute refined probability p^=1N∑j∈Ip′j\hat{p} = \frac{1}{N} \sum_{j \in I} p'^j and pseudo label y^=arg⁡max⁡cp^c\hat{y} = \arg\max_c \hat{p}^c
            Compute query feature q=ft(ts(xt))q = f_t(t_s(x_t)) and prediction pq=σ(ht(q))p_q = \sigma(h_t(q))
            Compute key feature k+=ft′(ts′(xt))k^+ = f'_t(t'_s(x_t))
            Compute momentum updates: w′=ft′(tw(xt))w' = f'_t(t_w(x_t)) and p′=σ(ht′(w′))p' = \sigma(h'_t(w'))
        end for
        
        Compute cross-entropy loss Ltce\mathcal{L}^{ce}_t between {pq}\{p_q\} and {y^}\{\hat{y}\}
        Compute contrastive loss Ltctr\mathcal{L}^{ctr}_t contrasting qq against k+k^+ and QsQ_s, excluding negatives where y^j=y^\hat{y}^j = \hat{y}
        Compute diversity loss Ltdiv\mathcal{L}^{div}_t on mean batch predictions pˉq\bar{p}_q
        Compute total loss Lt=Ltce+Ltctr+Ltdiv\mathcal{L}_t = \mathcal{L}^{ce}_t + \mathcal{L}^{ctr}_t + \mathcal{L}^{div}_t
        
        Update θt\theta_t via gradient descent on Lt\mathcal{L}_t
        Update momentum parameters: θt′←mθt′+(1−m)θt\theta'_t \leftarrow m\theta'_t + (1-m)\theta_t
        Enqueue batch features and probabilities {w′,p′}\{w', p'\} into QwQ_w, dequeue oldest
        Enqueue batch keys and pseudo labels {k+,y^}\{k^+, \hat{y}\} into QsQ_s, dequeue oldest
    end for
    return gtg_t
  6. Knowl 6 — Adaptation Performance on VisDA-C Benchmark

    data/table

    On the synthetic-to-real VisDA-C validation set (1212 classes), AdaContrast achieves state-of-the-art test-time adaptation performance. All methods use a ResNet-101 backbone except the On-target variants, which distill to a ResNet-18 student network.

    Method Source-free plane bcycl bus car horse knife mcycl person plant sktbrd train truck Avg.
    DANN no 81.9 77.7 82.8 44.3 81.2 29.5 65.1 28.6 51.9 54.6 82.8 7.8 57.4
    CDAN no 85.2 66.9 83.0 50.8 84.2 74.9 88.1 74.5 83.4 76.0 81.9 38.0 73.9
    CDAN+BSP no 92.4 61.0 81.0 57.5 89.0 80.6 90.1 77.0 84.2 77.9 82.1 38.4 75.9
    CAN no 97.0 87.2 82.5 74.3 97.8 96.2 90.8 80.7 96.6 96.3 87.5 59.9 87.2
    SWD no 90.8 82.5 81.7 70.5 91.7 69.5 86.3 77.5 87.4 63.6 85.6 29.2 76.4
    MCC no 88.7 80.3 80.5 71.5 90.1 93.2 85.0 71.6 89.4 73.8 85.0 36.9 78.8
    Source only - 57.2 11.1 42.4 66.9 55.0 4.4 81.1 27.3 57.9 29.4 86.7 5.8 43.8
    MA yes 94.8 73.4 68.8 74.8 93.1 95.4 88.6 84.7 89.1 84.7 83.5 48.1 81.6
    BAIT yes 93.7 83.2 84.5 65.0 92.9 95.4 88.1 80.8 90.0 89.0 84.0 45.3 82.7
    SHOT yes 95.3 87.5 78.7 55.6 94.1 94.2 81.4 80.0 91.8 90.7 86.5 59.8 83.0
    + On-target yes 96.0 89.5 84.3 67.2 95.9 94.2 91.0 81.5 93.8 89.9 89.1 58.2 85.9
    AdaContrast yes 97.0 84.7 84.0 77.3 96.7 93.8 91.9 84.8 94.3 93.1 94.1 49.7 86.8
    + On-target yes 97.2 87.0 86.7 81.7 95.5 91.6 93.5 86.6 95.3 90.9 92.8 47.9 87.2
    AdaContrast (online) yes 95.0 68.0 82.7 69.6 94.3 80.8 90.3 79.6 90.6 69.7 87.6 36.0 78.7

    AdaContrast reaches 86.8%86.8\% per-class average accuracy (84.5%84.5\% overall accuracy), outperforming SHOT by +3.8%+3.8\% average accuracy and +6.2%+6.2\% overall accuracy without accessing source data. When evaluated in a strict single-pass online streaming mode where pseudo label refinement is enabled after accumulating X=2048X = 2048 samples, AdaContrast achieves 78.7%78.7\% average accuracy.

  7. Knowl 7 — Adaptation Performance on DomainNet-126 Benchmark

    data/table

    AdaContrast was evaluated on DomainNet-126 across 7 domain shifts among 4 domains: Real (R), Clipart (C), Painting (P), and Sketch (S), using a ResNet-50 backbone.

    Method Source-free R→\toC R→\toP P→\toC C→\toS S→\toP R→\toS P→\toR Avg.
    MCC no 44.8 65.7 41.9 34.9 47.3 35.3 72.4 48.9
    Source only - 55.5 62.7 53.0 46.9 50.1 46.3 75.0 55.6
    TENT yes 58.5 65.7 57.9 48.5 52.4 54.0 67.0 57.7
    SHOT yes 67.7 68.4 66.9 60.1 66.1 59.9 80.8 67.1
    AdaContrast (Ours) yes 70.2 69.8 68.6 58.0 65.9 61.5 80.5 67.8
    AdaContrast (Ours, online) yes 61.1 66.9 60.8 53.4 62.7 54.5 78.9 62.6

    AdaContrast achieves 67.8%67.8\% average top-1 accuracy across the 7 domain shifts, outperforming the source-free baseline TENT (57.7%57.7\%) by +10.1%+10.1\%, SHOT (67.1%67.1\%) by +0.7%+0.7\%, and the source-present UDA baseline MCC (48.9%48.9\%) by +18.9%+18.9\%. In the online streaming adaptation setting, AdaContrast reaches 62.6%62.6\% average accuracy.

  8. Knowl 8 — Model Calibration Advantage over Entropy Minimization

    empirical result

    AdaContrast maintains superior target-domain model calibration compared to entropy minimization methods such as SHOT and TENT. Direct entropy minimization forces the model to output low-entropy, overly confident probability distributions on target data irrespective of label correctness, which distorts predictive uncertainty.

    On the VisDA-C validation set after adaptation:

    • SHOT achieves an Expected Calibration Error (ECE) of 2.97%2.97\% and a Maximum Calibration Error (MCE) of 39.16%39.16\%, displaying significant overconfidence on reliability diagrams.
    • AdaContrast achieves an ECE of 0.65%0.65\% and an MCE of 8.20%8.20\%, reducing ECE and MCE by more than a factor of 4.5×4.5\times relative to SHOT.

    This indicates that AdaContrast produces prediction probabilities that reflect true accuracy without sacrificing domain transfer performance.

  9. Knowl 9 — Ablation Study of AdaContrast Algorithmic Components

    data/table

    An ablation study isolates the contributions of online pseudo label refinement, joint contrastive learning with negative exclusion, and weak-strong/diversity regularization on DomainNet-126 (1×1\times learning rate) and VisDA-C (1×1\times and 10×10\times learning rates).

    # Pseudo labeling Online pl. ref Joint ctr. Reg. DN-126 (lr1x) VisDA-C (lr1x) VisDA-C (lr10x)
    0 55.6 43.8 43.8
    1 ✓ 58.5 55.0 44.5
    2 ✓ ✓ 64.7 86.5 9.9
    3 ✓ ✓ ✓ 67.9 85.7 84.3
    4 ✓ ✓ ✓ ✓ 67.8 86.8 86.6

    Key observations:

    1. Online pseudo-label refinement alone (row #2) boosts 1×1\times learning rate performance substantially (+6.2%+6.2\% on DN-126, +31.5%+31.5\% on VisDA-C) but diverges severely under a 10×10\times learning rate (9.9%9.9\% accuracy on VisDA-C) due to error compounding on noisy pseudo-labels without loss divergence.
    2. Enabling joint contrastive learning (row #3) stabilizes feature space representations, preventing pseudo-label drift and recovering VisDA-C 10×10\times learning rate accuracy from 9.9%9.9\% to 84.3%84.3\%.
    3. Excluding same-class negative pairs in the contrastive loss accounts for +0.2%+0.2\% gain on DN-126 (67.9%67.9\% vs 67.7%67.7\%) and +2.1%+2.1\% / +2.8%+2.8\% on VisDA-C (85.7%85.7\% vs 83.6%83.6\% at 1×1\times lr; 84.3%84.3\% vs 81.5%81.5\% at 10×10\times lr).
  10. Knowl 10 — Hyperparameter Insensitivity of AdaContrast

    empirical result

    AdaContrast exhibits insensitivity to its key hyperparameters:

    1. Learning Rate Scaling: When scaling the learning rate by 1×1\times, 3×3\times, and 10×10\times:

      • AdaContrast remains steady: DomainNet-126 averages are 67.8%67.8\%, 67.8%67.8\%, and 67.5%67.5\%; VisDA-C per-class averages are 86.8%86.8\%, 86.8%86.8\%, and 86.6%86.6\%.
      • SHOT drops significantly: DomainNet-126 averages fall from 67.1%67.1\% (1×1\times) to 64.7%64.7\% (10×10\times); VisDA-C per-class averages fall from 83.0%83.0\% (1×1\times) to 72.1%72.1\% (10×10\times).
    2. Memory Queue Size (MM): On VisDA-C, varying queue size across M∈{128,256,512,1024,2048,4096,8192,16384,32768,55388}M \in \{128, 256, 512, 1024, 2048, 4096, 8192, 16384, 32768, 55388\} yields stable performance. A queue size of M=512M = 512 achieves 84.4%84.4\% average accuracy, and M=2048M = 2048 (less than 4%4\% of the dataset) attains 86.3%86.3\%, on par with the full-dataset queue size (86.8%86.8\% at M=55388M=55388).

    3. Nearest Neighbors for Soft Voting (NN): Varying N∈{1,2,3,6,11,21,41}N \in \{1, 2, 3, 6, 11, 21, 41\} consistently achieves between 84.4%84.4\% and 86.8%86.8\% per-class average accuracy on VisDA-C.

Coverage note — None was omitted; all primary architectural components, mathematical formulas, algorithms, benchmark tables (VisDA-C and DomainNet-126), calibration statistics, and ablation analyses are fully covered.

References

  1. 1.Yuki Markus Asano, Christian Rupprecht, and Andrea Vedaldi. Self-labelling via simultaneous clustering and representation learning. arXiv preprint arXiv:1911.05371, 2019.
  2. 2.Konstantinos Bousmalis, Nathan Silberman, David Dohan, Dumitru Erhan, and Dilip Krishnan. Unsupervised pixel-level domain adaptation with generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3722–3731, 2017.
  3. 3.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In Proceedings of the European Conference on Computer Vision (ECCV), September 2018.
  4. 4.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  5. 5.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton. Big self-supervised models are strong semi-supervised learners. arXiv preprint arXiv:2006.10029, 2020.
  6. 6.Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020.
  7. 7.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15750–15758, 2021.
  8. 8.Xinyang Chen, Sinan Wang, Mingsheng Long, and Jianmin Wang. Transferability vs. discriminability: Batch spectral penalization for adversarial domain adaptation. In International conference on machine learning, pages 1081–1090. PMLR, 2019.
  9. 9.Yuhua Chen, Wen Li, Christos Sakaridis, Dengxin Dai, and Luc Van Gool. Domain adaptive faster r-cnn for object detection in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3339–3348, 2018.
  10. 10.Morris H DeGroot and Stephen E Fienberg. The comparison and evaluation of forecasters. Journal of the Royal Statistical Society: Series D (The Statistician), 32(1-2):12–22, 1983.
  11. 11.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  12. 12.Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In International conference on machine learning, pages 1180–1189. PMLR, 2015.
  13. 13.Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky. Domain-adversarial training of neural networks. The journal of machine learning research, 17(1):2096–2030, 2016.
  14. 14.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728, 2018.
  15. 15.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733, 2020.
  16. 16.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In International Conference on Machine Learning, pages 1321–1330. PMLR, 2017.
  17. 17.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9729–9738, 2020.
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  19. 19.Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei Efros, and Trevor Darrell. Cycada: Cycle-consistent adversarial domain adaptation. In International conference on machine learning, pages 1989–1998. PMLR, 2018.
  20. 20.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015.
  21. 21.Ying Jin, Ximei Wang, Mingsheng Long, and Jianmin Wang. Minimum class confusion for versatile domain adaptation. In European Conference on Computer Vision, pages 464–480. Springer, 2020.
  22. 22.Guoliang Kang, Lu Jiang, Yi Yang, and Alexander G Hauptmann. Contrastive adaptation network for unsupervised domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4893–4902, 2019.
  23. 23.Jogendra Nath Kundu, Naveen Venkat, R Venkatesh Babu, et al. Universal source-free domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4544–4553, 2020.
  24. 24.Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. Colorization as a proxy task for visual understanding. In CVPR, 2017.
  25. 25.Chen-Yu Lee, Tanmay Batra, Mohammad Haris Baig, and Daniel Ulbricht. Sliced wasserstein discrepancy for unsupervised domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10285–10295, 2019.
  26. 26.Dong-Hyun Lee et al. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML, volume 3, page 896, 2013.
  27. 27.Rui Li, Qianfen Jiao, Wenming Cao, Hau-San Wong, and Si Wu. Model adaptation: Unsupervised domain adaptation without source data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9641–9650, 2020.
  28. 28.Jian Liang, Dapeng Hu, and Jiashi Feng. Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation. In International Conference on Machine Learning, pages 6028–6039. PMLR, 2020.
  29. 29.Jian Liang, Dapeng Hu, and Jiashi Feng. Domain adaptation with auxiliary target domain-oriented classifier. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16632–16642, 2021.
  30. 30.Jian Liang, Dapeng Hu, Yunbo Wang, Ran He, and Jiashi Feng. Source data-absent unsupervised domain adaptation through hypothesis transfer and labeling transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  31. 31.Mingsheng Long, Yue Cao, Jianmin Wang, and Michael Jordan. Learning transferable features with deep adaptation networks. In International conference on machine learning, pages 97–105. PMLR, 2015.
  32. 32.Mingsheng Long, Zhangjie Cao, Jianmin Wang, and Michael I Jordan. Conditional adversarial domain adaptation. arXiv preprint arXiv:1705.10667, 2017.
  33. 33.HB Mitchell and PA Schaefer. A “soft” k-nearest neighbor voting scheme. International journal of intelligent systems, 16(4):459–468, 2001.
  34. 34.Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. Obtaining well calibrated probabilities using bayesian binning. In Twenty-Ninth AAAI Conference on Artificial Intelligence, 2015.
  35. 35.Alexandru Niculescu-Mizil and Rich Caruana. Predicting good probabilities with supervised learning. In Proceedings of the 22nd international conference on Machine learning, pages 625–632, 2005.
  36. 36.Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In European conference on computer vision, pages 69–84. Springer, 2016.
  37. 37.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  38. 38.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32:8026–8037, 2019.
  39. 39.Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for multi-source domain adaptation. In Proceedings of the IEEE International Conference on Computer Vision, pages 1406–1415, 2019.
  40. 40.Xingchao Peng, Ben Usman, Neela Kaushik, Judy Hoffman, Dequan Wang, and Kate Saenko. Visda: The visual domain adaptation challenge. arXiv preprint arXiv:1710.06924, 2017.
  41. 41.John Platt et al. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in large margin classifiers, 10(3):61–74, 1999.
  42. 42.Joaquin Quiñonero-Candela, Masashi Sugiyama, Neil D Lawrence, and Anton Schwaighofer. Dataset shift in machine learning. MIT Press, 2009.
  43. 43.Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and N Lawrence. Covariate shift and local learning by distribution matching, 2008.
  44. 44.Kate Saenko, Brian Kulis, Mario Fritz, and Trevor Darrell. Adapting visual category models to new domains. In European conference on computer vision, pages 213–226. Springer, 2010.
  45. 45.Kuniaki Saito, Donghyun Kim, Stan Sclaroff, Trevor Darrell, and Kate Saenko. Semi-supervised domain adaptation via minimax entropy. ICCV, 2019.
  46. 46.Kuniaki Saito, Donghyun Kim, Stan Sclaroff, and Kate Saenko. Universal domain adaptation through self supervision. arXiv preprint arXiv:2002.07953, 2020.
  47. 47.Tim Salimans and Durk P Kingma. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. Advances in neural information processing systems, 29:901–909, 2016.
  48. 48.Kihyuk Sohn, David Berthelot, Chun-Liang Li, Zizhao Zhang, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. arXiv preprint arXiv:2001.07685, 2020.
  49. 49.Yu Sun, Eric Tzeng, Trevor Darrell, and Alexei A Efros. Unsupervised domain adaptation through self-supervision. arXiv preprint arXiv:1909.11825, 2019.
  50. 50.Yu Sun, Xiaolong Wang, Liu Zhuang, John Miller, Moritz Hardt, and Alexei A. Efros. Test-time training with self-supervision for generalization under distribution shifts. In ICML, 2020.
  51. 51.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016.
  52. 52.Yi-Hsuan Tsai, Wei-Chih Hung, Samuel Schulter, Kihyuk Sohn, Ming-Hsuan Yang, and Manmohan Chandraker. Learning to adapt structured output space for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7472–7481, 2018.
  53. 53.Eric Tzeng, Judy Hoffman, Kate Saenko, and Trevor Darrell. Adversarial discriminative domain adaptation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7167–7176, 2017.
  54. 54.Eric Tzeng, Judy Hoffman, Ning Zhang, Kate Saenko, and Trevor Darrell. Deep domain confusion: Maximizing for domain invariance. arXiv preprint arXiv:1412.3474, 2014.
  55. 55.Dequan Wang, Shaoteng Liu, Sayna Ebrahimi, Evan Shelhamer, and Trevor Darrell. On-target adaptation. arXiv preprint arXiv:2109.01087, 2021.
  56. 56.Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. Tent: Fully test-time adaptation by entropy minimization. In International Conference on Learning Representations, 2021.
  57. 57.Shiqi Yang, Yaxing Wang, Joost van de Weijer, Luis Herranz, and Shangling Jui. Unsupervised domain adaptation without source data by casting a bait. arXiv preprint arXiv:2010.12427, 2020.
  58. 58.Shiqi Yang, Yaxing Wang, Joost van de Weijer, Luis Herranz, and Shangling Jui. Generalized source-free domain adaptation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8978–8987, 2021.
  59. 59.Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: Self-supervised learning via redundancy reduction. arXiv preprint arXiv:2103.03230, 2021.

Citation

MLA
Chen, D., et al. “Contrastive Test-Time Adaptation”. arXiv, 2022, http://arxiv.org/abs/2204.10377v1.
APA
Chen, D., Wang, D., Darrell, T., & Ebrahimi, S. (2022). Contrastive Test-Time Adaptation. arXiv. http://arxiv.org/abs/2204.10377v1
Chicago
Chen, D., D. Wang, T. Darrell, and S. Ebrahimi. 2022. “Contrastive Test-Time Adaptation”. arXiv. http://arxiv.org/abs/2204.10377v1.
Harvard
Chen, D. et al. (2022) “Contrastive Test-Time Adaptation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.10377v1.
Vancouver
1. Chen D, Wang D, Darrell T, Ebrahimi S (2022) Contrastive Test-Time Adaptation. arXiv

BibTeX

@article{chen2022contrastive,
  title = {Contrastive Test-Time Adaptation},
  author = {Chen, Dian and Wang, Dequan and Darrell, Trevor and Ebrahimi, Sayna},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.10377v1},
  eprint = {2204.10377}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE