Continual Test-Time Domain Adaptation

Qin WangOlga FinkLuc Van GoolDengxin Dai

article2022CVPR647 citations

Introduces CoTTA, a continual test-time adaptation framework that prevents error accumulation and catastrophic forgetting in dynamically shifting target domains through averaged predictions and stochastic weight restoration.

Listen

Real-world machine perception systems, such as autonomous vehicles, operate in non-stationary and continually changing environments where conditions like weather and lighting shift unpredictably. Adapting pre-trained neural networks to these shifts in real time is critical, but privacy constraints, bandwidth limits, or legal barriers frequently prevent access to the original source training data. Existing test-time adaptation methods assume stationary target domains and typically rely on self-training techniques that rapidly suffer from catastrophic forgetting and error accumulation when exposed to changing data streams.

The article develops and evaluates a continual test-time domain adaptation framework (CoTTA) designed to adapt off-the-shelf pre-trained models on unlabeled, streaming target data without requiring source data or model retraining.

The researchers assessed the approach across five continuous adaptation benchmarks, including image classification datasets with 15 corruption types and varying severities (CIFAR-10, CIFAR-100, and ImageNet) as well as an adverse-weather driving semantic segmentation task (Cityscapes to ACDC). The method addresses error propagation by generating refined self-training targets through an exponential moving average teacher network, applying confidence-triggered test-time data augmentations only during significant domain shifts. To counter catastrophic forgetting, the method stochastically restores approximately 1% of the model weights back to their original pre-trained values after each gradient update, allowing updates across all neural network layers rather than being restricted to batch normalization parameters.

The evaluation demonstrated that the proposed method consistently outperformed standard baselines across all tasks without destabilizing over long sequences. On the standard CIFAR-10 corruption benchmark, the method achieved an average classification error rate of 16.2%, significantly improving upon existing continuous entropy-minimization baselines (20.7%) and static models (43.5%). Under gradually shifting corruptions, the approach maintained an error rate of 10.4% compared to 30.7% for continuous entropy minimization. On the more challenging CIFAR-100 benchmark, it reduced the average error rate to 32.5% compared to 60.9% for continuous entropy baselines, which suffered severe performance degradation over time. On the adverse-condition driving segmentation task across repeated multi-round cycles, the approach achieved a 58.6% mean intersection-over-union score, preventing the long-term degradation observed in baseline models and maintaining stable performance across diverse neural network architectures, including vision transformers.

These findings indicate that effective continual adaptation does not require source data access or restricted parameter updates, provided that pseudo-label noise is controlled and source knowledge is periodically injected. In mission-critical deployments like autonomous transport and edge perception, this framework reduces operational safety risks and avoids the heavy infrastructure costs associated with continuously retraining models or transmitting centralized data.

Organizations deploying automated perception in dynamic operating environments should consider adopting weight-averaged pseudo-labeling alongside source weight restoration when deploying off-the-shelf models into the field. Prior to broader operational rollout, engineering teams should validate optimal confidence thresholds and evaluate computational overhead, as executing multiple test-time augmentations per frame introduces latency trade-offs that must be tuned to meet real-time processing requirements.

arXiv: 2203.13591
Cover for Continual Test-Time Domain Adaptation

Abstract

Test-time domain adaptation aims to adapt a source pre-trained model to a target domain without using any source data. Existing works mainly consider the case where the target domain is static. However, real-world machine perception systems are running in non-stationary and continually changing environments where the target domain distribution can change over time. Existing methods, which are mostly based on self-training and entropy regularization, can suffer from these non-stationary environments. Due to the distribution shift over time in the target domain, pseudo-labels become unreliable. The noisy pseudo-labels can further lead to error accumulation and catastrophic forgetting. To tackle these issues, we propose a continual test-time adaptation approach~(CoTTA) which comprises two parts. Firstly, we propose to reduce the error accumulation by using weight-averaged and augmentation-averaged predictions which are often more accurate. On the other hand, to avoid catastrophic forgetting, we propose to stochastically restore a small part of the neurons to the source pre-trained weights during each iteration to help preserve source knowledge in the long-term. The proposed method enables the long-term adaptation for all parameters in the network. CoTTA is easy to implement and can be readily incorporated in off-the-shelf pre-trained models. We demonstrate the effectiveness of our approach on four classification tasks and a segmentation task for continual test-time adaptation, on which we outperform existing methods. Our code is available at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Domain Adaptation
  • 2.2 Test-time Adaptation
  • 2.3 Continuous Domain Adaptation
  • 2.4 Continual Learning
  • 2.5 Domain Generalization
  • 3 Continual Test-Time Domain Adaptation
  • 3.1 Problem Definition
  • 3.2 Methodology
  • 4 Experiments
  • 4.1 Datasets and tasks
  • 4.2 Experiments on CIFAR10-to-CIFAR10C
  • 4.3 Experiments on CIFAR100-to-CIFAR100C
  • 4.4 Experiments on ImageNet-to-ImageNet-C
  • 4.5 Experiments on Cityscapes-to-ACDC
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Continual Test-Time Domain Adaptation Problem Formulation

    definition

    Online Continual Test-Time Domain Adaptation (CTTA) is the problem setting where an off-the-shelf neural network fθ0f_{\theta_0}, pre-trained on labeled source domain data (XS,YS)(\mathcal{X}^S, \mathcal{Y}^S) with parameters θ0\theta_0, is adapted online to an unlabeled target data stream XT\mathcal{X}^T without access to source data or future target data.

    At each discrete time step tt, the model receives a target input xtT∈XTx_t^T \in \mathcal{X}^T sampled from a non-stationary, continually shifting target distribution. The model must produce an online prediction y^tT=fθt(xtT)\hat{y}_t^T = f_{\theta_t}(x_t^T) immediately upon receiving xtTx_t^T, and subsequently update its parameters θt→θt+1\theta_t \to \theta_{t+1} before encountering xt+1Tx_{t+1}^T. Unlike standard unsupervised domain adaptation (which accesses labeled source data) or standard test-time adaptation (which assumes a single static target domain or allows offline batches/domain resets), CTTA operates under long-term non-stationary distribution shifts where the model must continually adapt without suffering from error accumulation or catastrophic forgetting.

  2. Knowl 2 — Teacher-Student Consistency via Exponential Moving Average Weights

    model/method

    To reduce error accumulation from noisy pseudo-labels during online continual adaptation, the Continual Test-Time Adaptation (CoTTA) framework uses a dual-network student-teacher architecture. The student network fθtf_{\theta_t} has trainable parameters θt\theta_t, and the teacher network fθt′f_{\theta'_t} maintains weight-averaged parameters θt′\theta'_t, both initialized to the source pre-trained model parameters θ0\theta_0.

    At time step tt, given target input xtTx_t^T, the student model generates class probability predictions y^tT=fθt(xtT)∈[0,1]C\hat{y}_t^T = f_{\theta_t}(x_t^T) \in [0, 1]^C across CC classes. The student is updated by minimizing the cross-entropy consistency loss with respect to teacher-derived soft pseudo-label targets yt′Ty'^T_t:

    Lθt(xtT)=−∑c=1Cytc′Tlog⁡y^tcT\mathcal{L}_{\theta_t}(x_t^T) = - \sum_{c=1}^C y'^T_{tc} \log \hat{y}^T_{tc}

    where ytc′Ty'^T_{tc} is the target probability assigned to class cc. After updating the student weights θt→θt+1\theta_t \to \theta_{t+1} via gradient descent, the teacher network weights are updated via an exponential moving average (EMA):

    θt+1′=αθt′+(1−α)θt+1\theta'_{t+1} = \alpha \theta'_t + (1 - \alpha) \theta_{t+1}

    where α∈[0,1)\alpha \in [0, 1) is a smoothing momentum factor. The final online prediction output for xtTx_t^T is the class with the highest probability under the teacher's prediction.

  3. Knowl 3 — Confidence-Gated Augmentation-Averaged Pseudo-Labels

    model/method

    To boost pseudo-label quality under severe domain distribution shifts while avoiding performance drops on familiar inputs, CoTTA estimates domain shift using the source model's confidence and applies test-time augmentation conditionally.

    Let conf(fθ0(xtT))=max⁡c(fθ0(xtT))c\text{conf}(f_{\theta_0}(x_t^T)) = \max_c (f_{\theta_0}(x_t^T))_c denote the prediction confidence of the original source pre-trained model fθ0f_{\theta_0} on the input xtTx_t^T, and let pthp_{th} be a predefined confidence threshold. The augmentation-averaged prediction from the teacher model across NN stochastic augmentations {augi}i=0N−1\{\text{aug}_i\}_{i=0}^{N-1} is computed as:

    y~t′T=1N∑i=0N−1fθt′(augi(xtT))\tilde{y}'^T_t = \frac{1}{N} \sum_{i=0}^{N-1} f_{\theta'_t}(\text{aug}_i(x_t^T))

    The refined teacher pseudo-label yt′Ty'^T_t supplied to the consistency loss is determined by:

    yt′T={y^t′T=fθt′(xtT)if conf(fθ0(xtT))≥pthy~t′Totherwisey'^T_t = \begin{cases} \hat{y}'^T_t = f_{\theta'_t}(x_t^T) & \text{if } \text{conf}(f_{\theta_0}(x_t^T)) \ge p_{th} \\ \tilde{y}'^T_t & \text{otherwise} \end{cases}

    When source confidence is high (indicating minimal domain shift), the unaugmented teacher prediction y^t′T\hat{y}'^T_t is used directly. When confidence falls below pthp_{th} (indicating a large domain gap), averaging predictions over NN augmented variants reduces pseudo-label noise and error propagation.

  4. Knowl 4 — Stochastic Parameter Restoration to Source Weights

    model/method

    To prevent catastrophic forgetting and avoid model collapse when updating all trainable layers over prolonged sequences of domain shifts, CoTTA applies stochastic restoration to the student network parameters after each gradient update.

    For any trainable weight tensor Wt+1W_{t+1} at step t+1t+1 (such as convolution or linear weight matrices) with corresponding source pre-trained tensor W0W_0, a binary mask tensor MM of identical shape is sampled element-wise from a Bernoulli distribution:

    M∼Bernoulli(p)M \sim \text{Bernoulli}(p)

    where p∈(0,1)p \in (0, 1) is a small restoration probability (e.g., p=0.01p = 0.01). The updated weight tensor is then computed by stochastic blending:

    Wt+1=M⊙W0+(1−M)⊙Wt+1W_{t+1} = M \odot W_0 + (1 - M) \odot W_{t+1}

    where ⊙\odot denotes the element-wise (Hadamard) product. By stochastically reverting a small fraction pp of weights back to their original source values W0W_0, the model prevents parameter drift away from the source distribution, acting as a structural regularizer that enables full-network continual adaptation rather than restricting updates solely to batch normalization parameters.

  5. Knowl 5 — CoTTA Online Continual Test-Time Adaptation Algorithm

    algorithm

    The Continual Test-Time Adaptation (CoTTA) algorithm adapts a pre-trained model online over a streaming sequence of unlabeled test samples from changing environments.

    Input: Source pre-trained model fθ0(x)f_{\theta_0}(x), data stream {xt}t=1T\{x_t\}_{t=1}^T, restore probability pp, smoothing factor α\alpha, confidence threshold pthp_{th}, number of augmentations NN
    Output: Online predictions {yt∗}t=1T\{y^*_t\}_{t=1}^T, adapted student fθT+1f_{\theta_{T+1}}, adapted teacher fθT+1′f_{\theta'_{T+1}}
    Initialize teacher network weights θ0′←θ0\theta'_0 \leftarrow \theta_0
    for each time step t=1,2,…,Tt = 1, 2, \dots, T do
        Receive target sample xtx_t
        Compute source confidence ct←max⁡c(fθ0(xt))cc_t \leftarrow \max_c (f_{\theta_0}(x_t))_c
        if ct≥pthc_t \ge p_{th} then
            yt′T←fθt′(xt)y'^T_t \leftarrow f_{\theta'_t}(x_t)
        else
            yt′T←1N∑i=0N−1fθt′(augi(xt))y'^T_t \leftarrow \frac{1}{N} \sum_{i=0}^{N-1} f_{\theta'_t}(\text{aug}_i(x_t))
        end if
        Compute online output prediction yt∗←arg⁡max⁡cytc′Ty^*_t \leftarrow \arg\max_c y'^T_{tc}
        Compute student prediction y^t←fθt(xt)\hat{y}_t \leftarrow f_{\theta_t}(x_t)
        Compute cross-entropy consistency loss L←−∑cytc′Tlog⁡y^tc\mathcal{L} \leftarrow - \sum_c y'^T_{tc} \log \hat{y}_{tc}
        Update student parameters θt+1←θt−η∇θtL\theta_{t+1} \leftarrow \theta_t - \eta \nabla_{\theta_t} \mathcal{L}
        Update teacher parameters θt+1′←αθt′+(1−α)θt+1\theta'_{t+1} \leftarrow \alpha \theta'_t + (1 - \alpha) \theta_{t+1}
        for each trainable tensor Wt+1∈θt+1W_{t+1} \in \theta_{t+1} with initial value W0W_0 do
            Sample mask M∼Bernoulli(p)M \sim \text{Bernoulli}(p)
            Wt+1←M⊙W0+(1−M)⊙Wt+1W_{t+1} \leftarrow M \odot W_0 + (1 - M) \odot W_{t+1}
        end for
    end for
  6. Knowl 6 — CIFAR-10 to CIFAR-10C Continual Adaptation and Component Ablation

    data/table

    Evaluation of online continual test-time domain adaptation on CIFAR-10 to CIFAR-10C at severity level 5 using WideResNet-28, with 10,000 images per corruption evaluated in continuous sequence across 15 corruption types. Hyperparameters: Adam optimizer (learning rate 10−310^{-3}), N=32N = 32 augmentations, p=0.01p = 0.01.

    Method W-avg A-avg S-Rest Gaus shot imp defoc glass motn zoom snow frost fog brit cntrst elast pixel jpeg Mean
    Source 72.3 65.7 72.9 46.9 54.3 34.8 42.0 25.1 41.3 26.0 9.3 46.7 26.6 58.5 30.3 43.5
    BN Stats Adapt 28.1 26.1 36.3 12.8 35.3 14.2 12.1 17.3 17.4 15.3 8.4 12.6 23.8 19.7 27.3 20.4
    Pseudo-label 26.7 22.1 32.0 13.8 32.2 15.3 12.7 17.3 17.3 16.5 10.1 13.4 22.4 18.9 25.9 19.8
    TENT-online* 24.8 23.5 33.0 12.0 31.8 13.7 10.8 15.9 16.2 13.7 7.9 12.1 22.0 17.3 24.2 18.6
    TENT-continual 24.8 20.6 28.6 14.4 31.1 16.5 14.1 19.1 18.6 18.6 12.2 20.3 25.7 20.8 24.9 20.7
    CoTTA (Ours) ✓ 27.2 22.8 30.8 12.1 30.1 13.9 11.9 17.2 16.0 14.3 9.4 13.1 19.9 15.4 19.9 18.3
    CoTTA (Ours) ✓ ✓ 24.5 21.0 26.0 12.3 27.9 13.9 12.0 16.6 15.9 14.7 9.4 13.6 19.8 14.7 18.7 17.4
    CoTTA (Ours) ✓ ✓ ✓ 24.3 21.3 26.6 11.6 27.6 12.2 10.3 14.8 14.1 12.4 7.5 10.6 18.3 13.4 17.3 16.2

    Note: TENT-online requires oracle domain boundary information to reset model weights at each new domain, which is unavailable in standard CTTA. TENT-continual exhibits severe error accumulation, degrading to 24.9% on the final domain (mean 20.7%). Each component of CoTTA incrementally reduces mean error: weight averaging (18.3%), augmentation averaging (17.4%), and stochastic restoration (16.2% with standard deviation 0.1% over 5 seeds).

  7. Knowl 7 — Gradually Changing Corruption Evaluation on CIFAR-10C

    data/table

    In realistic environments, distribution shifts occur gradually rather than in abrupt severity jumps. In the gradually changing setup, severity transitions smoothly through 1→2→3→4→5→4→3→2→11 \to 2 \to 3 \to 4 \to 5 \to 4 \to 3 \to 2 \to 1 within each corruption type, with domain transitions occurring at severity level 1. Results are averaged over 10 randomly shuffled corruption sequences.

    Metric Source Baseline BN Stats Adapt TENT-continual CoTTA (Ours)
    Avg. Error Rate (%) 24.8 13.7 30.7 10.4 ±\pm 0.2

    Under gradual shift, TENT-continual suffers catastrophic error accumulation and parameter drift, reaching 30.7% average error (worse than the non-adapted source baseline of 24.8%). CoTTA reduces the error rate to 10.4%, demonstrating robust stability under continuous transitions.

  8. Knowl 8 — Continual Test-Time Adaptation on CIFAR-100C

    data/table

    Classification error rate (%) on the 100-class CIFAR-100 to CIFAR-100C benchmark under continual online adaptation across 15 corruption types at severity level 5 using a pre-trained ResNeXt-29 architecture.

    Method Gaus shot imp defoc glass motn zoom snow frost fog brit cntrst elast pixel jpeg Mean
    Source 73.0 68.0 39.4 29.3 54.1 30.8 28.8 39.5 45.8 50.3 29.5 55.1 37.2 74.7 41.2 46.4
    BN Stats Adapt 42.1 40.7 42.7 27.6 41.9 29.7 27.9 34.9 35.0 41.5 26.5 30.3 35.7 32.9 41.2 35.4
    Pseudo-label 38.1 36.1 40.7 33.2 45.9 38.3 36.4 44.0 45.6 52.8 45.2 53.5 60.1 58.1 64.5 46.2
    TENT-continual 37.2 35.8 41.7 37.9 51.2 48.3 48.5 58.4 63.7 71.1 70.4 82.3 88.0 88.5 90.4 60.9
    CoTTA (Ours) 40.1 37.7 39.7 26.9 38.0 27.9 26.4 32.8 31.8 40.3 24.7 26.9 32.5 28.3 33.5 32.5

    TENT-continual collapses as corruptions progress (error rises from 37.2% on Gaussian to 90.4% on JPEG, with a mean of 60.9%). In contrast, CoTTA maintains stable adaptation throughout the sequence, achieving a mean error of 32.5% (an improvement of 2.9% over BN Stats Adapt and 13.9% over the unadapted source model).

  9. Knowl 9 — Continual Semantic Segmentation on Cityscapes-to-ACDC Benchmark

    data/table

    Long-term continual test-time semantic segmentation evaluated using a pre-trained Segformer-B5 model on Cityscapes adapted continually to four adverse conditions from ACDC (Fog →\to Night →\to Rain →\to Snow) repeated for 10 consecutive cycles (40 condition shifts total, 400 unlabeled images per condition). Predictions evaluated at original resolution (1920×10801920 \times 1080) using mean Intersection-over-Union (mIoU, %).

    Round 1 Round 4 Round 7 Round 10 All
    Method Fog Night Rain Snow Fog Night Rain Snow Fog Night Rain Snow Fog Night Rain Snow Mean
    Source 69.1 40.3 59.7 57.8 69.1 40.3 59.7 57.8 69.1 40.3 59.7 57.8 69.1 40.3 59.7 57.8 56.7
    BN Stats Adapt 62.3 38.0 54.6 53.0 62.3 38.0 54.6 53.0 62.3 38.0 54.6 53.0 62.3 38.0 54.6 53.0 52.0
    TENT-continual 69.0 40.2 60.1 57.3 66.5 36.3 58.7 54.0 64.2 32.8 55.3 50.9 61.8 29.8 51.9 47.8 52.3
    CoTTA (Ours) 70.9 41.2 62.4 59.7 70.9 41.0 62.7 59.7 70.9 41.0 62.8 59.7 70.8 41.0 62.8 59.7 58.6

    TENT-continual and BN Stats Adapt perform poorly on Segformer-B5 because Vision Transformers rely primarily on LayerNorm rather than BatchNorm; TENT's performance degrades steadily from 69.0% (Round 1 Fog) down to 29.8% (Round 10 Night) and 47.8% (Round 10 Snow). CoTTA adapts all trainable parameters without layer-type restrictions, maintaining consistent high performance (58.6% overall mIoU vs 56.7% source) across all 10 recurring cycles.

  10. Knowl 10 — Continual Test-Time Adaptation on ImageNet-to-ImageNet-C

    data/table

    Average classification error (%) on ImageNet-to-ImageNet-C evaluated on a pre-trained ResNet-50 across 10 diverse corruption sequences at corruption severity level 5.

    Metric Source Baseline BN Stats Adapt Test Aug TENT-continual CoTTA (Ours)
    Avg. Error Rate (%) 82.4 72.1 71.4 66.5 63.0 ±\pm 1.8 (0.1)

    Across 10 diverse sequence orderings of corruption types, CoTTA achieves an average error rate of 63.0±1.8%63.0 \pm 1.8\%, consistently outperforming the source baseline (82.4%), test-time batch statistics adaptation (72.1%), test-time augmentation baseline (71.4%), and TENT-continual (66.5%).

Coverage note — None was omitted; all key contributions including problem setup, algorithmic formulation, mathematical modeling, and empirical benchmarks on CIFAR-10C, CIFAR-100C, ImageNet-C, and Cityscapes-to-ACDC were included.

References

  1. 1.Arsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, and Dmitry Vetrov. Pitfalls of in-domain uncertainty estimation and ensembling in deep learning. ICLR, 2020. 3
  2. 2.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016. 8
  3. 3.Andreea Bobu, Eric Tzeng, Judy Hoffman, and Trevor Darrell. Adapting to continuously shifting domains. In ICLR Workshops, 2018. 3
  4. 4.Chaoqi Chen, Weiping Xie, Wenbing Huang, Yu Rong, Xinghao Ding, Yue Huang, Tingyang Xu, and Junzhou Huang. Progressive feature alignment for unsupervised domain adaptation. In CVPR, pages 627–636, 2019. 2
  5. 5.Gilad Cohen and Raja Giryes. Katana: Simple post-training robustness using test time augmentations. arXiv preprint arXiv:2109.08191, 2021. 5, 6
  6. 6.MMSegmentation Contributors. MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. https://github.com/open-mmlab/mmsegmentation, 2020. 6
  7. 7.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. 6
  8. 8.Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. Robustbench: a standardized adversarial robustness benchmark. In NeuIPS Datasets and Benchmarks Track, 2021. 6
  9. 9.Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le. Autoaugment: Learning augmentation strategies from data. In CVPR, pages 113–123, 2019. 5
  10. 10.Matthias Delange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Greg Slabaugh, and Tinne Tuytelaars. A continual learning survey: Defying forgetting in classification tasks. T-PAMI, 2021. 3
  11. 11.Sayna Ebrahimi, Franziska Meier, Roberto Calandra, Trevor Darrell, and Marcus Rohrbach. Adversarial continual learning. In ECCV. Springer, 2020. 2
  12. 12.Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In ICML, pages 1180–1189, 2015. 2
  13. 13.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In ICML, pages 1321–1330. PMLR, 2017. 2
  14. 14.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In ICCV, pages 8340–8349, 2021. 3
  15. 15.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. ICLR, 2019. 6
  16. 16.Dan Hendrycks, Norman Mu, Ekin D. Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. AugMix: A simple data processing method to improve robustness and uncertainty. ICLR, 2020. 3, 6
  17. 17.Judy Hoffman, Trevor Darrell, and Kate Saenko. Continuous manifold based adaptation for evolving visual domains. In CVPR, pages 867–874, 2014. 3
  18. 18.Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei Efros, and Trevor Darrell. Cycada: Cycle-consistent adversarial domain adaptation. In ICML, pages 1989–1998. PMLR, 2018. 2
  19. 19.Lukas Hoyer, Dengxin Dai, and Luc Van Gool. Daformer: Improving network architectures and training strategies for domain-adaptive semantic segmentation. arXiv preprint arXiv:2111.14887, 2021. 2
  20. 20.Minhao Hu, Tao Song, Yujun Gu, Xiangde Luo, Jieneng Chen, Yinan Chen, Ya Zhang, and Shaoting Zhang. Fully test-time adaptation for image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 251–260. Springer, 2021. 3
  21. 21.Xuefeng Hu, Gokhan Uzunbas, Sirius Chen, Rui Wang, Ashish Shah, Ram Nevatia, and Ser-Nam Lim. Mixnorm: Test-time adaptation through online normalization estimation. arXiv preprint arXiv:2110.11478, 2021. 3
  22. 22.Yusuke Iwasawa and Yutaka Matsuo. Test-time classifier adjustment module for model-agnostic domain generalization. In NeuIPS, 2021. 3
  23. 23.Neerav Karani, Ertunc Erdil, Krishna Chaitanya, and Ender Konukoglu. Test-time adaptable neural networks for robust medical image segmentation. Medical Image Analysis, 68:101907, 2021. 3
  24. 24.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 114(13):3521–3526, 2017. 3
  25. 25.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Technical report, Citeseer, 2009. 6
  26. 26.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In NeuIPS, pages 1097–1105, 2012. 5
  27. 27.Jogendra Nath Kundu, Akshay Kulkarni, Amit Singh, Varun Jampani, and R Venkatesh Babu. Generalize then adapt: Source-free domain adaptive semantic segmentation. In ICCV, pages 7046–7056, 2021. 3
  28. 28.Jogendra Nath Kundu, Naveen Venkat, R Venkatesh Babu, et al. Universal source-free domain adaptation. In CVPR, pages 4544–4553, 2020. 2
  29. 29.Vinod K Kurmi, Venkatesh K Subramanian, and Vinay P Namboodiri. Domain impression: A source data free domain adaptation method. In WACV, pages 615–625, 2021. 2
  30. 30.Qicheng Lao, Xiang Jiang, Mohammad Havaei, and Yoshua Bengio. Continuous domain adaptation with variational domain-agnostic feature replay. arXiv preprint arXiv:2003.04382, 2020. 3
  31. 31.Dong-Hyun Lee. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on Challenges in Representation Learning, ICML, volume 3, page 2, 2013. 7
  32. 32.Boyi Li, Felix Wu, Ser-Nam Lim, Serge Belongie, and Kilian Q Weinberger. On feature normalization and data augmentation. In CVPR, pages 12383–12392, 2021. 3
  33. 33.Rui Li, Qianfen Jiao, Wenming Cao, Hau-San Wong, and Si Wu. Model adaptation: Unsupervised domain adaptation without source data. In CVPR, pages 9641–9650, 2020. 2
  34. 34.Yanghao Li, Naiyan Wang, Jianping Shi, Jiaying Liu, and Xiaodi Hou. Revisiting batch normalization for practical domain adaptation. arXiv preprint arXiv:1603.04779, 2016. 3, 7
  35. 35.Zhizhong Li and Derek Hoiem. Learning without forgetting. T-PAMI, 40(12):2935–2947, 2017. 3
  36. 36.Qing Lian, Fengmao Lv, Lixin Duan, and Boqing Gong. Constructing self-motivated pyramid curriculums for cross-domain semantic segmentation: A non-adversarial approach. In ICCV, October 2019. 2
  37. 37.Jian Liang, Dapeng Hu, and Jiashi Feng. Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation. In ICLR, pages 6028–6039. PMLR, 2020. 2
  38. 38.Yuang Liu, Wei Zhang, and Jun Wang. Source-free domain adaptation for semantic segmentation. In CVPR, pages 1215–1224, 2021. 3
  39. 39.Mingsheng Long, Yue Cao, Jianmin Wang, and Michael Jordan. Learning transferable features with deep adaptation networks. In ICML, pages 97–105, 2015. 2
  40. 40.Alexander Lyzhov, Yuliya Molchanova, Arsenii Ashukha, Dmitry Molchanov, and Dmitry Vetrov. Greedy policy search: A simple baseline for learnable test-time augmentation. In Jonas Peters and David Sontag, editors, Conference on Uncertainty in Artificial Intelligence (UAI), volume 124 of Proceedings of Machine Learning Research, pages 1308–1317. PMLR, 03–06 Aug 2020. 3
  41. 41.Michael McCloskey and Neal J Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of learning and motivation, volume 24, pages 109–165. Elsevier, 1989. 2
  42. 42.Krikamol Muandet, David Balduzzi, and Bernhard Schölkopf. Domain generalization via invariant feature representation. In ICML, pages 10–18. PMLR, 2013. 3
  43. 43.Chaithanya Kumar Mummadi, Robin Hutmacher, Kilian Rambach, Evgeny Levinkov, Thomas Brox, and Jan Hendrik Metzen. Test-time adaptation to distribution shift by confidence maximization and input transformation. arXiv preprint arXiv:2106.14999, 2021. 1, 2, 3, 6
  44. 44.Sinno Jialin Pan, Ivor W Tsang, James T Kwok, and Qiang Yang. Domain adaptation via transfer component analysis. IEEE Transactions on Neural Networks, 22(2):199–210, 2011. 2
  45. 45.German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter. Continual lifelong learning with neural networks: A review. Neural Networks, 113:54–71, 2019. 2, 3
  46. 46.Vishal M Patel, Raghuraman Gopalan, Ruonan Li, and Rama Chellappa. Visual domain adaptation: A survey of recent advances. IEEE signal processing magazine, 32(3):53–69, 2015. 2
  47. 47.Boris T Polyak and Anatoli B Juditsky. Acceleration of stochastic approximation by averaging. SIAM journal on control and optimization, 30(4):838–855, 1992. 4, 5
  48. 48.Viraj Prabhu, Shivam Khare, Deeksha Kartik, and Judy Hoffman. Sentry: Selective entropy optimization via committee consistency for unsupervised domain adaptation. In ICCV, pages 8558–8567, 2021. 2
  49. 49.Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert. icarl: Incremental classifier and representation learning. In CVPR, pages 2001–2010, 2017. 3
  50. 50.Christos Sakaridis, Dengxin Dai, and Luc Van Gool. ACDC: The adverse conditions dataset with correspondences for semantic driving scene understanding. In ICCV, October 2021. 1, 6
  51. 51.Steffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann, Wieland Brendel, and Matthias Bethge. Improving robustness against common corruptions by covariate shift adaptation. NeuIPS, 33, 2020. 7
  52. 52.Connor Shorten and Taghi M Khoshgoftaar. A survey on image data augmentation for deep learning. Journal of Big Data, 6(1):1–48, 2019. 3, 5
  53. 53.Daniel L Silver and Robert E Mercer. The task rehearsal method of life-long learning: Overcoming impoverished data. In Conference of the Canadian Society for Computational Studies of Intelligence, pages 90–101. Springer, 2002. 3
  54. 54.Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, and Moritz Hardt. Test-time training with self-supervision for generalization under distribution shifts. In ICML, pages 9229–9248. PMLR, 2020. 3, 4, 5
  55. 55.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In NeuIPS, pages 1195–1204, 2017. 2, 4, 5
  56. 56.Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In IROS, pages 23–30. IEEE, 2017. 3
  57. 57.Jonathan Tremblay, Aayush Prakash, David Acuna, Mark Brophy, Varun Jampani, Cem Anil, Thang To, Eric Cameracci, Shaad Boochoon, and Stan Birchfield. Training deep networks with synthetic data: Bridging the reality gap by domain randomization. In CVPR Workshops, pages 969–977, 2018. 3
  58. 58.Y.-H. Tsai, W.-C. Hung, S. Schulter, K. Sohn, M.-H. Yang, and M. Chandraker. Learning to adapt structured output space for semantic segmentation. In CVPR, 2018. 2
  59. 59.Riccardo Volpi, Diane Larlus, and Grégory Rogez. Continual adaptation of visual representations via domain randomization and meta-learning. In CVPR, pages 4443–4453, 2021. 3
  60. 60.Dequan Wang, Shaoteng Liu, Sayna Ebrahimi, Evan Shelhamer, and Trevor Darrell. On-target adaptation. arXiv preprint arXiv:2109.01087, 2021. 3
  61. 61.Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. Tent: Fully test-time adaptation by entropy minimization. In ICLR, 2021. 1, 2, 3, 4, 6, 7, 8
  62. 62.Qin Wang, Dengxin Dai, Lukas Hoyer, Luc Van Gool, and Olga Fink. Domain adaptive semantic segmentation with self-supervised depth estimation. In ICCV, pages 8515–8525, 2021. 2
  63. 63.Markus Wulfmeier, Alex Bewley, and Ingmar Posner. Incremental adversarial domain adaptation for continually changing environments. In ICRA, pages 4489–4495. IEEE, 2018. 3
  64. 64.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. In NeuIPS, 2021. 6
  65. 65.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In CVPR, pages 1492–1500, 2017. 6
  66. 66.Shiqi Yang, Yaxing Wang, Joost van de Weijer, Luis Herranz, and Shangling Jui. Generalized source-free domain adaptation. In ICCV, pages 8978–8987, 2021. 2
  67. 67.Yanchao Yang and Stefano Soatto. Fda: Fourier domain adaptation for semantic segmentation. In CVPR, pages 4085–4095, 2020. 2
  68. 68.Hao-Wei Yeh, Baoyao Yang, Pong C Yuen, and Tatsuya Harada. Sofa: Source-data-free feature alignment for unsupervised domain adaptation. In WACV, pages 474–483, 2021. 2
  69. 69.Dong Yin, Raphael Gontijo Lopes, Jonathon Shlens, Ekin D Cubuk, and Justin Gilmer. A fourier perspective on model robustness in computer vision. NeuIPS, 2019. 3
  70. 70.Fuming You, Jingjing Li, and Zhou Zhao. Test-time batch statistics calibration for covariate shift. arXiv preprint arXiv:2110.04065, 2021. 3
  71. 71.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. BMVC, 2016. 6
  72. 72.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In International Conference on Machine Learning, pages 3987–3995. PMLR, 2017. 3
  73. 73.Marvin Zhang, Sergey Levine, and Chelsea Finn. Memo: Test time robustness via adaptation and augmentation. arXiv preprint arXiv:2110.09506, 2021. 3
  74. 74.Aurick Zhou and Sergey Levine. Training on test data with bayesian adaptation for covariate shift. arXiv preprint arXiv:2109.12746, 2021. 3
  75. 75.Yang Zou, Zhiding Yu, BVK Vijaya Kumar, and Jinsong Wang. Unsupervised domain adaptation for semantic segmentation via class-balanced self-training. In ECCV, pages 289–305, 2018. 2

Citation

MLA
Wang, Q., et al. “Continual Test-Time Domain Adaptation”. arXiv, 2022, http://arxiv.org/abs/2203.13591v1.
APA
Wang, Q., Fink, O., Gool, L. V., & Dai, D. (2022). Continual Test-Time Domain Adaptation. arXiv. http://arxiv.org/abs/2203.13591v1
Chicago
Wang, Q., O. Fink, L. V. Gool, and D. Dai. 2022. “Continual Test-Time Domain Adaptation”. arXiv. http://arxiv.org/abs/2203.13591v1.
Harvard
Wang, Q. et al. (2022) “Continual Test-Time Domain Adaptation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.13591v1.
Vancouver
1. Wang Q, Fink O, Gool LV, Dai D (2022) Continual Test-Time Domain Adaptation. arXiv

BibTeX

@article{wang2022continual,
  title = {Continual Test-Time Domain Adaptation},
  author = {Wang, Qin and Fink, Olga and Gool, Luc Van and Dai, Dengxin},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.13591v1},
  eprint = {2203.13591}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE