Feature Alignment and Uniformity for Test Time Adaptation

Shuai WangDaoan ZhangZipei YanJianguo ZhangRui Li

article2023CVPR74 citations

Proposes a test-time adaptation framework that simultaneously optimizes feature alignment and uniformity via online self-distillation, memorized local clustering, and noise-filtering mechanisms to adapt pre-trained models to out-of-distribution data.

Listen

Deploying machine learning models in production often leads to significant performance drops when real-world data differs from the data used during training. Addressing these data distribution shifts typically requires costly retraining or access to labeled target data, both of which are impractical in dynamic, real-time operating environments. Test time adaptation addresses this challenge by adjusting a pre-trained model directly during inference using only incoming, unlabeled data streams.

The article introduces a framework that models test time adaptation as a dual-objective feature revision problem, simultaneously optimizing feature uniformity and feature alignment. The approach operates entirely online: it establishes uniformity through unsupervised test time self-distillation linked to historical data stored in a memory bank, while maintaining alignment through memorized spatial local clustering among neighboring data points. Dual filtering mechanisms—specifically entropy and prediction consistency filters—are embedded to identify and discard unreliable, noisy pseudo-labels during online updates.

Empirical evaluations across four standard domain generalization benchmarks and four cross-domain medical image segmentation tasks demonstrate substantial performance improvements. On classification benchmarks, the method improved baseline accuracy by up to 4.8 percentage points on standard architectures and surpassed existing state-of-the-art test time adaptation and source-free adaptation techniques. In medical segmentation tasks, the framework consistently outperformed base models across prostate, cardiac, and retinal datasets, increasing baseline segmentation accuracy by up to 12.7 percentage points.

These findings indicate that addressing representation quality directly during deployment can reliably mitigate domain shifts without modifying the initial training process or requiring source data access. For decision-makers, this provides a practical, plug-and-play solution to improve AI reliability and safety in high-stakes fields such as healthcare. Resource analysis shows manageable operational overhead: reducing test batch sizes or updating only batch normalization parameters reduces graphics processing memory by approximately 40% to 50% with negligible loss in accuracy.

Organizations deploying visual AI models under variable operating conditions should evaluate and pilot this test time adaptation framework to improve operational robustness. Teams should tune batch sizes and parameter updates to match their specific hardware constraints. Additional investigation is recommended before deploying the method to low-level image processing tasks, such as denoising or super-resolution, where discrete category definitions do not apply.

arXiv: 2303.10902
Cover for Feature Alignment and Uniformity for Test Time Adaptation

Abstract

Test time adaptation (TTA) aims to adapt deep neural networks when receiving out of distribution test domain samples. In this setting, the model can only access online unlabeled test samples and pre-trained models on the training domains. We first address TTA as a feature revision problem due to the domain gap between source domains and target domains. After that, we follow the two measurements alignment and uniformity to discuss the test time feature revision. For test time feature uniformity, we propose a test time self-distillation strategy to guarantee the consistency of uniformity between representations of the current batch and all the previous batches. For test time feature alignment, we propose a memorized spatial local clustering strategy to align the representations among the neighborhood samples for the upcoming batch. To deal with the common noisy label problem, we propound the entropy and consistency filters to select and drop the possible noisy labels. To prove the scalability and efficacy of our method, we conduct experiments on four domain generalization benchmarks and four medical image segmentation tasks with various backbones. Experiment results show that our method not only improves baseline stably but also outperforms existing state-of-the-art test time adaptation methods.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Domain Generalization
  • 2.2. Test Time Adaptation
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Test Time Self-Distillation
  • 3.3. Memorized Spatial Local Clustering
  • 3.4. Training Objective Function
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Comparative Study
  • 4.3. Ablation Study
  • 4.4. Analysis
  • 4.5. Scalability to Medical Image Segmentation
  • 5. Limitations
  • 6. Conclusion
  • Acknowledgment
  • References

Knowls

  1. Knowl 1 — Test Time Self-Distillation with Entropy and Consistency Filtering

    model/method

    Test Time Self-Distillation (TSD) optimizes test-time feature uniformity in an online adaptation setting where target domain samples arrive sequentially. The model g=f∘hg = f \circ h consists of a feature backbone ff and a linear classification head hh producing logits pi=h(f(xi))∈RCp_i = h(f(x_i)) \in \mathbb{R}^C for an input xix_i, with predicted class y^i=arg⁡max⁡pi\hat{y}_i = \arg\max p_i.

    A memory bank B={(zi,pi)}\mathcal{B} = \{(z_i, p_i)\} stores feature embeddings zi=f(xi)z_i = f(x_i) and logits pip_i of all arriving samples, initialized with the weights of the source classifier head hh. To ensure feature uniformity between current samples and past data, class prototypes ckc_k are calculated as the mean embedding of samples assigned to class kk:

    ck=∑iziI[y^i=k]∑iI[y^i=k]c_k = \frac{\sum_i z_i \mathbb{I}[\hat{y}_i = k]}{\sum_i \mathbb{I}[\hat{y}_i = k]}

    where I[⋅]\mathbb{I}[\cdot] is the indicator function.

    To prevent incorrect pseudo-labels from corrupting prototypes, an Entropy Filter (EF) computes the Shannon entropy for each prediction:

    H(pi)=−∑k=1Cσ(pi)klog⁡σ(pi)kH(p_i) = -\sum_{k=1}^C \sigma(p_i)_k \log \sigma(p_i)_k

    where σ(⋅)\sigma(\cdot) denotes the softmax function. Embeddings with the top-MM highest entropy values for each class in the memory bank are excluded from prototype computation.

    The prototype-based classification distribution for sample ii is calculated via cosine similarity sim(zi,ck)=zi⊤ck∥zi∥∥ck∥\text{sim}(z_i, c_k) = \frac{z_i^\top c_k}{\|z_i\| \|c_k\|}:

    yik=exp⁡(sim(zi,ck))∑k′=1Cexp⁡(sim(zi,ck′))y_i^k = \frac{\exp(\text{sim}(z_i, c_k))}{\sum_{k'=1}^C \exp(\text{sim}(z_i, c_{k'}))}

    The self-distillation loss minimizes the divergence between the model's soft pseudo-label distribution σ(pi)\sigma(p_i) and the prototype distribution yiy_i:

    Li(pi,yi)=−σ(pi)log⁡yi=−∑k=1Cσ(pi)klog⁡yik\mathcal{L}_i(p_i, y_i) = -\sigma(p_i) \log y_i = -\sum_{k=1}^C \sigma(p_i)_k \log y_i^k

    To suppress gradient updates from noisy samples, a Consistency Filter (CF) generates a binary reliability mask MiM_i indicating whether the linear classifier and the prototype classifier predict the same top class:

    Mi=I[arg⁡max⁡pi=arg⁡max⁡yi]M_i = \mathbb{I}[\arg\max p_i = \arg\max y_i]

    The overall TSD loss over reliable samples in a batch is:

    Ltsd=∑iLi(pi,yi)⋅Mi∑iMi\mathcal{L}_{\text{tsd}} = \frac{\sum_i \mathcal{L}_i(p_i, y_i) \cdot M_i}{\sum_i M_i}

  2. Knowl 2 — Memorized Spatial Local Clustering for Test Time Feature Alignment

    model/method

    Memorized Spatial Local Clustering (MSLC) enforces feature alignment during online test-time adaptation by constraining the representation of an incoming target sample to be consistent with its local neighborhood in the latent space.

    For a target image xx with embedding z=f(x)z = f(x) and logit output p=h(z)p = h(z), the KK-nearest neighbor feature embeddings {zj}j=1K\{z_j\}_{j=1}^K and their corresponding logits {pj}j=1K\{p_j\}_{j=1}^K are retrieved from the memory bank B\mathcal{B}. The prediction distribution of xx is regularized toward the prediction distributions of its spatial neighbors, weighted by cosine similarity sim(z,zj)=z⊤zj∥z∥∥zj∥\text{sim}(z, z_j) = \frac{z^\top z_j}{\|z\| \|z_j\|}:

    Lmslc=1K∑j=1Ksim(z,zj)(σ(p)−σ(pj))2\mathcal{L}_{\text{mslc}} = \frac{1}{K} \sum_{j=1}^K \text{sim}(z, z_j) (\sigma(p) - \sigma(p_j))^2

    where σ(⋅)\sigma(\cdot) denotes the softmax operator.

    To avoid a trivial collapsed solution where all target representations map to a constant output regardless of the input, the gradient of the similarity weighting term sim(z,zj)\text{sim}(z, z_j) is detached during backpropagation and treated as a constant scalar.

    The overall test-time adaptation objective combines self-distillation and spatial local clustering:

    L=Ltsd+λLmslc\mathcal{L} = \mathcal{L}_{\text{tsd}} + \lambda \mathcal{L}_{\text{mslc}}

    where λ\lambda is a trade-off parameter (set empirically to λ=0.1\lambda = 0.1 by default).

  3. Knowl 3 — Online Test-Time Adaptation via Alignment and Uniformity

    algorithm

    The online test-time adaptation algorithm adapts a model g=f∘hg = f \circ h pre-trained on source domain data DsD_s to an unlabelled target stream {xt}t=1N\{x_t\}_{t=1}^N without accessing source data or running multiple training loops.

    Input: Pre-trained source model g=f∘hg = f \circ h, unlabeled test stream {xt}t=1N\{x_t\}_{t=1}^N, batch size BB, trade-off λ=0.1\lambda = 0.1, nearest neighbors KK, entropy threshold MM, learning rate η\eta
    Output: Target predictions {pt}t=1N\{p_t\}_{t=1}^N
    Initialize memory bank B={(wk,wk)}k=1C\mathcal{B} = \{(w_k, w_k)\}_{k=1}^C using classifier head weights
    for each incoming batch {xi}i=1B\{x_i\}_{i=1}^B do
        Extract embeddings zi=f(xi)z_i = f(x_i) and logits pi=h(zi)p_i = h(z_i)
        Insert {(zi,pi)}i=1B\{(z_i, p_i)\}_{i=1}^B into memory bank B\mathcal{B}
        Compute class prototypes ckc_k from B\mathcal{B}, excluding the top-MM highest-entropy samples per class
        for each sample i∈{1,…,B}i \in \{1, \dots, B\} do
            Compute prototype-based distribution yik=exp⁡(sim(zi,ck))∑k′=1Cexp⁡(sim(zi,ck′))y_i^k = \frac{\exp(\text{sim}(z_i, c_k))}{\sum_{k'=1}^C \exp(\text{sim}(z_i, c_{k'}))}
            Compute consistency mask Mi=I[arg⁡max⁡pi=arg⁡max⁡yi]M_i = \mathbb{I}[\arg\max p_i = \arg\max y_i]
            Retrieve KK-nearest neighbor embeddings {zj}j=1K\{z_j\}_{j=1}^K and logits {pj}j=1K\{p_j\}_{j=1}^K of ziz_i from B\mathcal{B}
        end for
        Compute TSD loss Ltsd=∑i−σ(pi)log⁡(yi)⋅Mi∑iMi\mathcal{L}_{\text{tsd}} = \frac{\sum_i -\sigma(p_i) \log(y_i) \cdot M_i}{\sum_i M_i}
        Compute MSLC loss Lmslc=1B∑i=1B1K∑j=1Ksim(zi,zj)(σ(pi)−σ(pj))2\mathcal{L}_{\text{mslc}} = \frac{1}{B} \sum_{i=1}^B \frac{1}{K} \sum_{j=1}^K \text{sim}(z_i, z_j) (\sigma(p_i) - \sigma(p_j))^2 (with sim(zi,zj)\text{sim}(z_i, z_j) detached)
        Compute total loss L=Ltsd+λLmslc\mathcal{L} = \mathcal{L}_{\text{tsd}} + \lambda \mathcal{L}_{\text{mslc}}
        Perform one-step gradient descent update on all trainable model parameters: θ←θ−η∇θL\theta \leftarrow \theta - \eta \nabla_\theta \mathcal{L}
    end for
  4. Knowl 4 — Domain Generalization Benchmark Classification Performance

    data/table

    Test-time adaptation accuracy (%) across four standard domain generalization benchmarks (PACS, OfficeHome, VLCS, and DomainNet) evaluated using ResNet-18 and ResNet-50 backbones. The reported results are the averages of three independent runs with varying weight initializations, random seeds, and data splits.

    Method PACS OfficeHome VLCS DomainNet Avg.
    ResNet-18 Backbone
    ERM 82.07 63.12 72.75 38.95 64.22
    BN 82.82 62.30 64.31 37.80 61.81
    Tent 84.92 63.75 67.36 38.95 63.75
    PL 84.64 60.22 68.93 35.23 62.26
    SHOT-IM 82.55 63.42 64.90 39.50 62.59
    T3A 83.50 64.25 73.03 39.61 65.10
    ETA 82.70 62.46 64.35 39.43 62.24
    LAME 84.58 62.20 72.88 37.49 64.29
    Ours 87.32 64.83 73.61 40.19 66.49
    ResNet-50 Backbone
    ERM 84.59 67.37 74.01 45.20 67.74
    BN 85.03 66.10 64.78 43.38 64.82
    Tent 87.48 67.96 69.20 44.71 67.34
    PL 85.23 67.13 68.52 41.18 65.52
    SHOT-IM 85.50 67.39 65.23 46.30 66.11
    T3A 86.04 68.29 73.98 46.16 68.62
    ETA 85.04 66.21 64.79 46.13 65.54
    LAME 86.62 66.19 73.94 43.20 67.49
    Ours 89.41 68.67 74.52 47.73 70.08

    The combined alignment and uniformity strategy outperforms both source models (ERM) and existing TTA methods across all benchmarks, improving average accuracy by +2.27% on ResNet-18 and +2.34% on ResNet-50 over ERM.

  5. Knowl 5 — Comparison with Domain Generalization and Source-Free Domain Adaptation Methods

    data/table

    Comparison of test-time adaptation performance against state-of-the-art Domain Generalization (DG) methods and Source-Free Domain Adaptation (SFDA) methods using a ResNet-50 backbone.

    On the PACS benchmark (domains: Art (A), Cartoon (C), Photo (P), Sketch (S)):

    Method A C P S Avg.
    ERM 82.5 80.8 94.1 81.0 84.6
    DNA 89.8 83.4 97.7 82.6 88.4
    PCL 90.2 83.9 98.1 82.6 88.7
    SWAD 89.3 83.2 96.9 83.4 88.2
    Ours 87.7 88.8 96.2 85.0 89.4
    SWAD + Ours 92.2 89.2 97.1 85.6 91.0

    On the DomainNet benchmark (6 domains: clip, info, paint, quick, real, sketch):

    Method clip info paint quick real sketch Avg.
    Domain Generalization Methods
    ERM 64.8 22.1 51.8 13.8 64.7 54.0 45.2
    PCL 67.9 24.3 55.3 15.7 66.6 56.4 47.7
    DNA 66.1 23.0 54.6 16.7 65.8 56.8 47.2
    SWAD 66.1 22.4 53.6 16.3 65.5 56.2 46.7
    Source-Free Domain Adaptation Methods
    F-mix 75.4 24.6 57.8 23.6 65.8 58.5 51.0
    Test-Time Adaptation (Ours)
    Ours 66.1 24.1 52.8 18.2 68.5 56.7 47.7
    SWAD + Ours 69.2 28.4 58.2 26.2 68.1 59.6 51.6

    Applying test-time adaptation on top of SWAD-trained weights achieves the top average performance on both PACS (91.0%) and DomainNet (51.6%), exceeding offline SFDA methods (F-mix) while requiring only single-pass online test stream processing.

  6. Knowl 6 — Ablation and Hyperparameter Sensitivity Analysis

    data/table

    Ablation study on PACS using ResNet-50 evaluating the individual contributions of Self-Distillation (SD), Entropy Filter (EF), Consistency Filter (CF), and Memorized Spatial Local Clustering (MSLC).

    # SD EF CF MSLC Acc (% ↑\uparrow)
    0 84.59
    1 ✓ 87.80 (+3.21)
    2 ✓ ✓ 88.40 (+3.81)
    3 ✓ ✓ ✓ 88.92 (+4.33)
    4 ✓ ✓ ✓ ✓ 89.41 (+4.82)

    Hyperparameter sensitivity findings on PACS:

    • Nearest neighbors KK in MSLC: Best performance is attained for K∈{1,3,5}K \in \{1, 3, 5\} (peak accuracy ≈89.41%\approx 89.41\% at K=3K=3). Increasing KK to {10,15,20}\{10, 15, 20\} causes degradation (dropping towards ≈87.5%\approx 87.5\%) due to the incorporation of misclassified neighboring features.
    • Entropy filter parameter MM: Performance is stable across M∈{1,5,20,50,100}M \in \{1, 5, 20, 50, 100\}, yielding monotonic gains with peak performance near M=50–100M=50\text{--}100. Omitting EF (NA) drops accuracy slightly to ≈88.9%\approx 88.9\%.
    • Trade-off parameter λ\lambda: Robust within λ∈[0.1,0.5]\lambda \in [0.1, 0.5], with λ=0.1\lambda = 0.1 delivering the best accuracy.
  7. Knowl 7 — Backbone Scalability Across Diverse Architectures

    empirical result

    The alignment and uniformity test-time adaptation framework operates in a model-agnostic, plug-and-play manner across convolutional, transformer, and MLP architectures on PACS, OfficeHome, and VLCS:

    Backbone PACS OfficeHome VLCS
    ResNet18 82.07 63.12 72.75
    + Ours 87.32 (+5.25) 64.83 (+1.71) 73.61 (+0.86)
    ResNet50 84.59 67.37 74.01
    + Ours 89.41 (+4.82) 68.67 (+1.30) 74.52 (+0.51)
    ResNeXt-50 86.67 72.66 78.50
    + Ours 91.33 (+4.66) 74.18 (+1.52) 79.38 (+0.88)
    ViT-B/16 87.13 79.06 78.70
    + Ours 90.20 (+3.07) 81.80 (+2.74) 79.90 (+1.20)
    EfficientNet-B4 85.11 74.65 77.14
    + Ours 85.41 (+0.30) 72.24 (-2.41) 79.42 (+2.28)
    Mixer-L16 84.59 71.36 76.53
    + Ours 88.47 (+3.88) 74.82 (+3.46) 79.75 (+3.22)

    The adaptation strategy consistently enhances performance across backbones without modifying network structures or requiring architecture-specific layer selections.

  8. Knowl 8 — Cross-Domain Medical Image Segmentation Performance

    data/table

    Evaluation of test-time adaptation on four cross-domain medical image segmentation benchmarks using DeepLabv3+ with MobileNetV2 backbone. Evaluated using average Dice Score (%):

    1. Prostate Segmentation: Source dataset Promise12 (50 cases) →\to Target dataset MSD05 (32 cases).
    2. Cardiac Structure Segmentation: Source dataset ACDC (200 cases) →\to Target dataset LGE Multi-sequence Cardiac MR (45 cases), segmenting LV blood pool, RV blood pool, and myocardium.
    3. Optic Disc and Cup Segmentation: Source dataset REFUGE (400 images) →\to Target dataset Drishti-GS (101 images).
    4. Optic Disc and Cup Segmentation: Source dataset REFUGE (400 images) →\to Target dataset RIM-ONE-r3 (159 images).
    Source Promise12 ACDC REFUGE
    Target MSD05 LGE DrishtiGS RIM-ONE-r3
    ERM 83.00 82.34 81.05 68.17
    Tent 85.64 81.11 83.96 77.87
    Ours 86.90 86.78 84.62 80.89

    For segmentation tasks, KK is set to 64 in MSLC. The method improves ERM by +3.90% on MSD05, +4.44% on LGE, +3.57% on DrishtiGS, and +12.72% on RIM-ONE-r3.

  9. Knowl 9 — Computational Efficiency and Batch Normalization Parameter Adaptation

    empirical result

    Computational cost and memory scaling during test-time adaptation depend primarily on the test batch size:

    • Batch Size Scaling: On PACS with ResNet-50, the default batch size of 128 achieves 89.4% accuracy with ≈14 GB\approx 14\text{ GB} GPU memory usage. Reducing batch size to 64 achieves 89.1% accuracy while reducing GPU memory consumption to 7.85 GB7.85\text{ GB} and yielding the fastest wall-clock execution time.
    • Affine Parameter Fine-Tuning: Rather than updating all model parameters, adapting only the affine parameters (scale and shift) of Batch Normalization layers achieves 89.0%89.0\% accuracy on PACS (a minor decrease of 0.4% from the 89.4%89.4\% achieved when tuning all parameters) while saving 2 GB2\text{ GB} of GPU memory.
  10. Knowl 10 — Limitations in Low-Level Vision and Test-Time Hyperparameter Selection

    limitation

    The feature alignment and uniformity framework has two key limitations:

    1. Inapplicability to Low-Level Vision Tasks: The formulation relies fundamentally on semantic class prototypes, category predictions, and Shannon prediction entropy. In low-level pixel tasks such as image dehazing, denoising, and super-resolution, semantic prototypes and categorical entropy do not exist, preventing direct application.
    2. Discrepancy in Hyperparameter Tuning: All hyperparameters (KK, MM, λ\lambda, learning rate) must be tuned prior to accessing test data using training/source domain validation sets. Optimal hyperparameter values identified on source domains may not coincide with the optimal configurations for arbitrary target distributions under severe domain shift.

Coverage note — No substantial contributed material was omitted. All methodological components (TSD, EF, CF, MSLC), algorithmic workflow, benchmark classification results, SFDA/DG comparisons, backbone scalability, medical segmentation experiments, ablations, complexity analyses, and limitations are fully covered.

References

  1. 1.Michela Antonelli, Annika Reinke, Spyridon Bakas, Keyvan Farahani, et al. The medical segmentation decathlon. Nature Communications, 13(1), July 2022. 2, 8
  2. 2.Yogesh Balaji, Swami Sankaranarayanan, and Rama Chellappa. Metareg: Towards domain generalization using meta-regularization. In NeurIPS, 2018. 3
  3. 3.Mikhail Belkin and Partha Niyogi. Laplacian eigenmaps for dimensionality reduction and data representation. Neural Computation, 15(6):1373–1396, June 2003. 4
  4. 4.Olivier Bernard, Alain Lalande, Clement Zotti, et al. Deep learning techniques for automatic MRI cardiac multi-structures segmentation and diagnosis: Is the problem solved? IEEE Transactions on Medical Imaging, 37(11):2514–2525, Nov. 2018. 2, 8
  5. 5.Malik Boudiaf, Romain Muller, Ismail Ben Ayed, and Luca Bertinetto. Parameter-free online test-time adaptation. In CVPR, 2022. 1, 3, 5
  6. 6.Fabio Maria Carlucci, Antonio D’Innocente, Silvia Bucci, Barbara Caputo, and Tatiana Tommasi. Domain generalization by solving jigsaw puzzles. In CVPR, 2019. 3
  7. 7.Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park. SWAD: domain generalization by seeking flat minima. In NeurIPS, 2021. 6
  8. 8.Dian Chen, Dequan Wang, Trevor Darrell, and Sayna Ebrahimi. Contrastive test-time adaptation. In CVPR, 2022. 1, 3
  9. 9.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018. 8
  10. 10.Xu Chu, Yujie Jin, Wenwu Zhu, Yasha Wang, Xin Wang, Shanghang Zhang, and Hong Mei. DNA: domain generalization with diversified neural averaging. In ICML, 2022. 6
  11. 11.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 5, 7
  12. 12.Qi Dou, Daniel Coelho de Castro, Konstantinos Kamnitsas, and Ben Glocker. Domain generalization via model-agnostic learning of semantic features. In NeurIPS, 2019. 3
  13. 13.F. Fumero, S. Alayon, J. L. Sanchez, J. Sigut, and M. Gonzalez-Hernandez. RIM-ONE: An open retinal image database for optic nerve evaluation. In 2011 24th International Symposium on Computer-Based Medical Systems (CBMS). IEEE, June 2011. 2, 8
  14. 14.Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario March, and Victor Lempitsky. Domain-adversarial training of neural networks. Journal of Machine Learning Research, 17(59):1–35, 2016. 3
  15. 15.Spyros Gidaris, Praveer Singh, and Nikos Komodakis. Unsupervised representation learning by predicting image rotations. In ICLR, 2018. 1, 3
  16. 16.Ishaan Gulrajani and David Lopez-Paz. In search of lost domain generalization. In ICLR, 2021. 5
  17. 17.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 1, 5, 7
  18. 18.Dan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. Augmix: A simple data processing method to improve robustness and uncertainty. In ICLR, 2020. 3
  19. 19.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 4
  20. 20.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015. 3, 5
  21. 21.Yusuke Iwasawa and Yutaka Matsuo. Test-time classifier adjustment module for model-agnostic domain generalization. In NeurIPS, 2021. 1, 2, 3, 5
  22. 22.Daehee Kim, Youngjun Yoo, Seunghyun Park, Jinkyu Kim, and Jaekoo Lee. Selfreg: Self-supervised contrastive regularization for domain generalization. In ICCV, 2021. 3
  23. 23.Kyungyul Kim, Byeongmoon Ji, Doyoung Yoon, and Sangheum Hwang. Self-knowledge distillation with progressive refinement of targets. In ICCV, 2021. 2
  24. 24.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 5, 8
  25. 25.Takeshi Kojima, Yutaka Matsuo, and Yusuke Iwasawa. Robustifying vision transformer without retraining from scratch by test-time class-conditional feature alignment. In IJCAI, 2022. 1, 3
  26. 26.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In NeurIPS, 2012. 1
  27. 27.Jogendra Nath Kundu, Akshay R. Kulkarni, Suvaansh Bhambri, Deepesh Mehta, Shreyas Anand Kulkarni, Varun Jampani, and Venkatesh Babu Radhakrishnan. Balancing discriminability and transferability for source-free domain adaptation. In ICML, 2022. 6
  28. 28.Dong-Hyun Lee et al. Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on challenges in representation learning, ICML, 2013. 5
  29. 29.Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M. Hospedales. Learning to generalize: Meta-learning for domain generalization. In AAAI, 2018. 3
  30. 30.Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M Hospedales. Deeper, broader and artier domain generalization. In ICCV, 2017. 2, 5, 8
  31. 31.Haoliang Li, Sinno Jialin Pan, Shiqi Wang, and Alex C. Kot. Domain generalization with adversarial feature learning. In CVPR, 2018. 3
  32. 32.Xiaotong Li, Yongxing Dai, Yixiao Ge, Jun Liu, Ying Shan, and Lingyu Duan. Uncertainty modeling for out-of-distribution generalization. In ICLR, 2022. 3
  33. 33.Yanghao Li, Naiyan Wang, Jianping Shi, Xiaodi Hou, and Jiaying Liu. Adaptive batch normalization for practical domain adaptation. Pattern Recognition, 80:109–117, Aug. 2018. 2, 3
  34. 34.Jian Liang, Dapeng Hu, and Jiashi Feng. Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation. In ICML, 2020. 2, 3, 5
  35. 35.Geert Litjens, Robert Toth, Wendy van de Ven, Caroline Hoeks, Sjoerd Kerkstra, Bram van Ginneken, Graham Vincent, Gwenael Guillard, Neil Birbeck, Jindang Zhang, et al. Evaluation of prostate segmentation algorithms for mri: the promise12 challenge. Medical image analysis, 18(2):359–373, 2014. 2, 8
  36. 36.Yuejiang Liu, Parth Kothari, Bastien van Delft, Baptiste Bellot-Gurlet, Taylor Mordan, and Alexandre Alahi. TTT++: when does self-supervised test-time training fail or thrive? In NeurIPS, 2021. 1, 3
  37. 37.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 1
  38. 38.Fangrui Lv, Jian Liang, Shuang Li, Bin Zang, Chi Harold Liu, Ziteng Wang, and Di Liu. Causality inspired representation learning for domain generalization. In CVPR, 2022. 3
  39. 39.Divyat Mahajan, Shruti Tople, and Amit Sharma. Domain generalization using causal matching. In ICML, 2021. 3
  40. 40.Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 3DV, 2016. 8
  41. 41.Saeid Motiian, Marco Piccirilli, Donald A. Adjeroh, and Gianfranco Doretto. Unified deep supervised domain adaptation and generalization. In ICCV, 2017. 3
  42. 42.Zachary Nado, Shreyas Padhy, D Sculley, Alexander D’Amour, Balaji Lakshminarayanan, and Jasper Snoek. Evaluating prediction-time batch normalization for robustness under covariate shift. arXiv preprint arXiv:2006.10963, 2020. 2, 3
  43. 43.Shuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen, Shijian Zheng, Peilin Zhao, and Mingkui Tan. Efficient test-time model adaptation without forgetting. In ICML, 2022. 1, 2, 3, 5
  44. 44.Oren Nuriel, Sagie Benaim, and Lior Wolf. Permuted adain: Reducing the bias towards global statistics in image classification. In CVPR, 2021. 3
  45. 45.José Ignacio Orlando, Huazhu Fu, João Barbosa Breda, et al. REFUGE challenge: A unified framework for evaluating automated methods for glaucoma assessment from fundus photographs. Medical Image Analysis, 59:101570, Jan. 2020. 2, 8
  46. 46.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019. 5
  47. 47.Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for multi-source domain adaptation. In ICCV, 2019. 2, 5
  48. 48.Alexandre Rame, Corentin Dancette, and Matthieu Cord. Fishr: Invariant gradient variances for out-of-distribution generalization. In International Conference on Machine Learning, pages 18347–18377. PMLR, 2022. 8
  49. 49.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, Apr. 2015. 5
  50. 50.Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, 2018. 8
  51. 51.Steffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann, Wieland Brendel, and Matthias Bethge. Improving robustness against common corruptions by covariate shift adaptation. In NeurIPS, 2020. 2, 3, 5
  52. 52.Claude E. Shannon. A mathematical theory of communication. Bell System Technical Journal, 27(3):379–423, 1948. 3
  53. 53.Jayanthi Sivaswamy, S Krishnadas, Arunava Chakravarty, G Joshi, A Syed Tabish, et al. A comprehensive retinal image dataset for the assessment of glaucoma from the optic nerve head analysis. JSM Biomedical Imaging Data Papers, 2(1):1004, 2015. 2, 8
  54. 54.Baochen Sun and Kate Saenko. Deep CORAL: correlation alignment for deep domain adaptation. In ECCV, 2016. 3
  55. 55.Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, and Moritz Hardt. Test-time training with self-supervision for generalization under distribution shifts. In ICML, 2020. 1, 3
  56. 56.Mingxing Tan and Quoc V. Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, 2019. 5, 7
  57. 57.Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. Mlp-mixer: An all-mlp architecture for vision. In NeurIPS, 2021. 5, 7
  58. 58.Antonio Torralba and Alexei A Efros. Unbiased look at dataset bias. In CVPR, 2011. 2, 5
  59. 59.Eric Tzeng, Judy Hoffman, Ning Zhang, Kate Saenko, and Trevor Darrell. Deep domain confusion: Maximizing for domain invariance. arXiv preprint arXiv:1412.3474, 2014. 3
  60. 60.Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9(86):2579–2605, 2008. 7
  61. 61.Vladimir Vapnik. Statistical learning theory. Wiley, 1998. 5, 6, 7, 8
  62. 62.Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty, and Sethuraman Panchanathan. Deep hashing network for unsupervised domain adaptation. In CVPR, 2017. 2, 5
  63. 63.Vikas Verma, Alex Lamb, Christopher Beckham, Amir Najafi, Ioannis Mitliagkas, David Lopez-Paz, and Yoshua Bengio. Manifold mixup: Better representations by interpolating hidden states. In ICML, 2019. 3
  64. 64.Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. Tent: Fully test-time adaptation by entropy minimization. In ICLR, 2021. 1, 2, 3, 5, 8
  65. 65.Qin Wang, Olga Fink, Luc Van Gool, and Dengxin Dai. Continual test-time domain adaptation. In CVPR, 2022. 1, 3
  66. 66.Tongzhou Wang and Phillip Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In ICML, 2020. 1
  67. 67.Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In CVPR, 2017. 5, 7
  68. 68.Xufeng Yao, Yang Bai, Xinyun Zhang, Yuechen Zhang, Qi Sun, Ran Chen, Ruiyu Li, and Bei Yu. PCL: proxy-based contrastive learning for domain generalization. In CVPR, 2022. 3, 6
  69. 69.Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh, Youngjoon Yoo, and Junsuk Choe. Cutmix: Regularization strategy to train strong classifiers with localizable features. In ICCV, 2019. 3
  70. 70.Daoan Zhang, Mingkai Chen, Chenming Li, Lingyun Huang, and Jianguo Zhang. Aggregation of disentanglement: Reconsidering domain variations in domain generalization. arXiv preprint arXiv:2302.02350, 2023. 3
  71. 71.Daoan Zhang, Chenming Li, Haoquan Li, Wenjian Huang, Lingyun Huang, and Jianguo Zhang. Rethinking alignment and uniformity in unsupervised image semantic segmentation. arXiv preprint arXiv:2211.14513, 2022. 1
  72. 72.Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2018. 3
  73. 73.Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma. Be your own teacher: Improve the performance of convolutional neural networks via self distillation. In ICCV, 2019. 2
  74. 74.Marvin Mengxin Zhang, Sergey Levine, and Chelsea Finn. MEMO: Test time robustness via adaptation and augmentation. In Advances in Neural Information Processing Systems, 2022. 1, 2
  75. 75.Kaiyang Zhou, Yongxin Yang, Yu Qiao, and Tao Xiang. Domain generalization with mixstyle. In ICLR, 2021. 3
  76. 76.Xiahai Zhuang. Multivariate mixture model for cardiac segmentation from multi-sequence MRI. In MICCAI. 2016. 2, 8
  77. 77.Xiahai Zhuang. Multivariate mixture model for myocardial segmentation combining multi-source images. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(12):2933–2946, Dec. 2019. 2, 8

Citation

MLA
Wang, S., et al. “Feature Alignment and Uniformity for Test Time Adaptation”. arXiv, 2023, http://arxiv.org/abs/2303.10902v3.
APA
Wang, S., Zhang, D., Yan, Z., Zhang, J., & Li, R. (2023). Feature Alignment and Uniformity for Test Time Adaptation. arXiv. http://arxiv.org/abs/2303.10902v3
Chicago
Wang, S., D. Zhang, Z. Yan, J. Zhang, and R. Li. 2023. “Feature Alignment and Uniformity for Test Time Adaptation”. arXiv. http://arxiv.org/abs/2303.10902v3.
Harvard
Wang, S. et al. (2023) “Feature Alignment and Uniformity for Test Time Adaptation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.10902v3.
Vancouver
1. Wang S, Zhang D, Yan Z, Zhang J, Li R (2023) Feature Alignment and Uniformity for Test Time Adaptation. arXiv

BibTeX

@article{wang2023feature,
  title = {Feature Alignment and Uniformity for Test Time Adaptation},
  author = {Wang, Shuai and Zhang, Daoan and Yan, Zipei and Zhang, Jianguo and Li, Rui},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.10902v3},
  eprint = {2303.10902}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE