SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models

Xiaosong MaJie ZhangSong GuoWenchao Xu

article2023NeurIPS75 citations

Proposes SwapPrompt, a test-time prompt adaptation framework that pairs an exponential moving average prompt with a self-supervised swapped prediction mechanism to boost the accuracy of frozen vision-language models on unlabeled target domains without requiring backbone fine-tuning.

Listen

Deploying artificial intelligence in real-world settings often exposes a fundamental vulnerability: when incoming test data differs from original training data, predictive accuracy declines significantly. Traditional solutions require access to source datasets or computationally expensive fine-tuning of massive model backbones, which is frequently impractical during live deployment. While emerging pre-trained vision-language models provide strong foundational representations, adapting them using existing unlabeled test-time methods often leads to over-confidence risks or suboptimal accuracy.

The article demonstrates a novel framework called SwapPrompt, which adapts vision-language models to new, unlabeled data environments at test time without modifying the underlying model backbone. Its objective is to evaluate how combining self-supervised contrastive learning with dynamic prompt adaptation improves domain generalization across diverse image classification tasks.

To achieve this, the approach freezes the core neural network and focuses exclusively on optimizing lightweight text prompts. It introduces a dual-prompt architecture where an actively updated online prompt learns alongside a slowly moving target prompt that preserves historical information. By comparing different augmented views of the same unlabeled image against dynamic class prototypes, the system uses a swapped prediction mechanism combined with high-confidence pseudo-label optimization. The authors validated this methodology across 14 benchmark datasets representing standard image recognition, natural distribution shifts, and fine-grained classification tasks.

The evaluation produced several key findings. First, SwapPrompt achieves state-of-the-art test-time adaptation, outperforming leading unsupervised baselines—such as Test-Time Prompt Tuning and Unsupervised Prompt Learning—by an average margin of 2.31% and 2.17% across all 14 datasets. Second, the framework reaches competitive performance with supervised few-shot adaptation methods and even surpasses them on select benchmarks despite having no access to ground-truth labels. Third, the system operates effectively in online data streaming scenarios with only minor drops in accuracy. Finally, efficiency analyses reveal rapid convergence, matching or exceeding prior baselines within just two adaptation cycles.

These results demonstrate that organizations can achieve robust, high-performance model adaptation at low computational cost. By adjusting only a handful of prompt parameters rather than retrained backbones, engineering teams can significantly reduce infrastructure expenses, lower operational latency, and mitigate deployment risks in shifting data environments. This challenges the assumption that expensive supervised annotations or full model retraining are necessary to handle distribution shifts.

For implementation, organizations operating large vision-language models should adopt contrastive prompt adaptation strategies for edge and real-time inference pipelines. Where inference speed is paramount, operators can limit adaptation to two or three cycles to capture the bulk of performance gains. To maintain adaptation quality, practitioners must filter incoming test streams to select only high-confidence samples, as uncurated pseudo-labels introduce performance-degrading noise.

Confidence in these findings is supported by extensive ablation studies and testing across varied visual domains. However, users should note that the framework's effectiveness remains tied to the quality of the underlying pre-trained model and can experience minor accuracy declines on certain specialized adversarial datasets. Further validation on non-image modalities and broader enterprise architectures is recommended before wide-scale deployment.

Ma et al (2023).pdf
  • Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). Introduces Context Optimization (CoOp) for learning continuous prompts on vision-language models like CLIP, establishing the foundational prompt adaptation formulation that SwapPrompt directly builds upon.
  • Paper: Conditional Prompt Learning for Vision-Language Models, Kaiyang Zhou et al. (2022). Extends prompt learning to input-conditional contexts (CoCoOp) to address generalization under domain shifts, providing essential context for SwapPrompt's test-time prompt adaptation setting.
  • Paper: Contrastive Test-Time Adaptation, Dian Chen et al. (2022). Establishes self-supervised contrastive learning and online pseudo-label refinement for test-time adaptation, foundational concepts adapted by SwapPrompt for run-time prompt tuning.
  • Paper: Continual Test-Time Domain Adaptation, Qin Wang et al. (2022). Demonstrates the use of exponential moving average target models and test-time augmentations for robust adaptation under domain shift, underpinning SwapPrompt's dual-prompt paradigm.
  • Paper: Test-Time Training with Self-Supervision for Generalization under Distribution Shifts, Yu Sun et al. (2019). Introduces the test-time training framework using self-supervised objectives on unlabeled test streams, laying the conceptual groundwork for adapting models during inference.
  • Paper: CLIP-Adapter: Better Vision-Language Models with Feature Adapters, Peng Gao et al. (2021). Presents parameter-efficient adaptation of vision-language models without fine-tuning full backbones, motivating SwapPrompt's focus on lightweight prompt-based updates.
  • Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). Introduces visual prompt tuning for transformer backbones, setting the stage for visual and multi-modal prompt tuning paradigms in vision-language models.
Cover for SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models

Abstract

Test-time adaptation (TTA) is a special and practical setting in unsupervised domain adaptation, which allows a pre-trained model in a source domain to adapt to unlabeled test data in another target domain. To avoid the computation-intensive backbone fine-tuning process, the zero-shot generalization potentials of the emerging pre-trained vision-language models (e.g., CLIP, CoOp) are leveraged to only tune the run-time prompt for unseen test domains. However, existing solutions have yet to fully exploit the representation capabilities of pre-trained models as they only focus on the entropy-based optimization and the performance is far below the supervised prompt adaptation methods, e.g., CoOp. In this paper, we propose SwapPrompt, a novel framework that can effectively leverage the self-supervised contrastive learning to facilitate the test-time prompt adaptation. SwapPrompt employs a dual prompts paradigm, i.e., an online prompt and a target prompt that averaged from the online prompt to retain historical information. In addition, SwapPrompt applies a swapped prediction mechanism, which takes advantage of the representation capabilities of pre-trained models to enhance the online prompt via contrastive learning. Specifically, we use the online prompt together with an augmented view of the input image to predict the class assignment generated by the target prompt together with an alternative augmented view of the same image. The proposed SwapPrompt can be easily deployed on vision-language models without additional requirement, and experimental results show that it achieves state-of-the-art test-time adaptation performance on ImageNet and nine other datasets. It is also shown that SwapPrompt can even achieve comparable performance with supervised prompt adaptation methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Preliminary and Problem Definition
  • 3.2 Overview of SwapPrompt
  • 3.2.1 Exponential Moving Average Prompt
  • 3.2.2 Prompt Swapped Prediction
  • 3.2.3 Prompt Optimization by Pseudo Label
  • 3.3 Algorithm Workflow
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Performance Comparison
  • 4.3 Ablation Study
  • 5 Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Test-time prompt adaptation problem for CLIP

    definition

    SwapPrompt addresses test-time prompt adaptation for a frozen pre-trained vision-language model. The model has an image encoder ff and text encoder gg; only a learnable continuous prompt tt is adapted using an unlabeled target-domain test set Dtest={xi}i=1N\mathcal{D}_{\mathrm{test}}=\{x_i\}_{i=1}^{N}, where each image belongs to one of CC classes but its label is unavailable. For class cc, the text feature is wc=g({t;c})w_c=g(\{t;c\}), and the image feature is zi=f(xi)z_i=f(x_i). The class probability is

    p(c∣xi;t)=exp⁡(sim⁡(zi,wc)/τ)∑j=1Cexp⁡(sim⁡(zi,wj)/τ),p(c\mid x_i;t)=\frac{\exp(\operatorname{sim}(z_i,w_c)/\tau)}{\sum_{j=1}^{C}\exp(\operatorname{sim}(z_i,w_j)/\tau)},

    where sim⁡\operatorname{sim} is cosine similarity and τ>0\tau>0 is the temperature. The adaptation objective is to learn a prompt t∗t^* from unlabeled test images while keeping the image and text encoders fixed.

  2. Knowl 2 — Dual-prompt exponential moving average mechanism

    model/method

    SwapPrompt maintains an online prompt tot_o and a target prompt ttt_t. The online prompt is optimized by the adaptation losses, whereas the target prompt is not directly optimized by back-propagation. After each online-prompt update, the target prompt is replaced by an exponential moving average of its previous value and the current online prompt:

    ttnew=ϵttold+(1−ϵ)to,t_t^{\mathrm{new}}=\epsilon t_t^{\mathrm{old}}+(1-\epsilon)t_o,

    where the decay rate satisfies ϵ∈[0,1]\epsilon\in[0,1]. A large ϵ\epsilon retains more historical information and makes the target prompt change slowly. The target prompt supplies delayed, more stable text representations that act as class prototypes, while the online prompt learns to predict the target prompt’s assignments. In the reported experiments, ϵ=0.99\epsilon=0.99.

  3. Knowl 3 — Prompt-swapped prediction loss

    model/method

    For an unlabeled image xix_i, SwapPrompt generates two augmented views A1(xi)A_1(x_i) and A2(xi)A_2(x_i). Let zir=f(Ar(xi))z_i^r=f(A_r(x_i)) for r∈{1,2}r\in\{1,2\}, let wct=g({tt;c})w_c^t=g(\{t_t;c\}) be the target-prompt text feature for class cc, and let wco=g({to;c})w_c^o=g(\{t_o;c\}) be the corresponding online-prompt text feature. The target-prompt class assignment for view rr is

    qir(c)=exp⁡(sim⁡(zir,wct)/τ)∑j=1Cexp⁡(sim⁡(zir,wjt)/τ),q_i^r(c)=\frac{\exp(\operatorname{sim}(z_i^r,w_c^t)/\tau)}{\sum_{j=1}^{C}\exp(\operatorname{sim}(z_i^r,w_j^t)/\tau)},

    and the online-prompt prediction is

    pir(c)=exp⁡(sim⁡(zir,wco)/τ)∑j=1Cexp⁡(sim⁡(zir,wjo)/τ).p_i^r(c)=\frac{\exp(\operatorname{sim}(z_i^r,w_c^o)/\tau)}{\sum_{j=1}^{C}\exp(\operatorname{sim}(z_i^r,w_j^o)/\tau)}.

    The swapped prediction loss trains the online prompt to predict the target-prompt assignment of the opposite view:

    Lswap(xi)=−∑c=1C[qi2(c)log⁡pi1(c)+qi1(c)log⁡pi2(c)].\mathcal{L}_{\mathrm{swap}}(x_i)=-\sum_{c=1}^{C}\left[q_i^2(c)\log p_i^1(c)+q_i^1(c)\log p_i^2(c)\right].

    The target assignments are treated as fixed targets for this loss, and the target prompt is updated only through the exponential moving average. The mechanism exploits the assumption that two augmentations of the same image should receive similar class assignments, while coupling the online prompt to a temporally stable target prompt.

  4. Knowl 4 — Pseudo-label objective and confidence-based data selection

    model/method

    Because target-domain labels are unavailable, SwapPrompt first applies zero-shot CLIP with a hand-crafted prompt such as “a photo of a [CLS]” to obtain a pseudo-label y^i\hat y_i and confidence for each test image. To reduce noisy pseudo-labels, it retains only the top KK highest-confidence images within each class, forming an adaptation subset Dadapt⊆Dtest\mathcal{D}_{\mathrm{adapt}}\subseteq\mathcal{D}_{\mathrm{test}}. For a selected image, the pseudo-label loss uses both augmented views:

    Lpseudo(xi)=ℓce(pi1,y^i)+ℓce(pi2,y^i),\mathcal{L}_{\mathrm{pseudo}}(x_i)=\ell_{\mathrm{ce}}(p_i^1,\hat y_i)+\ell_{\mathrm{ce}}(p_i^2,\hat y_i),

    where pi1p_i^1 and pi2p_i^2 are the online-prompt class distributions and ℓce\ell_{\mathrm{ce}} is cross-entropy with the class-index pseudo-label y^i\hat y_i. The total prompt-adaptation loss is

    Ladapt(xi)=αLswap(xi)+βLpseudo(xi),\mathcal{L}_{\mathrm{adapt}}(x_i)=\alpha\mathcal{L}_{\mathrm{swap}}(x_i)+\beta\mathcal{L}_{\mathrm{pseudo}}(x_i),

    where α\alpha and β\beta weight self-supervised swapped prediction and pseudo-label supervision. The standard setting uses α=β=1\alpha=\beta=1.

  5. Knowl 5 — Offline and online SwapPrompt workflow

    algorithm

    Input: a frozen vision-language model with image encoder ff, text encoder gg, an unlabeled target-domain test set or a stream of test mini-batches, class count CC, confidence-selection parameter KK, augmentation operators A1,A2A_1,A_2, and loss weights α,β\alpha,\beta.

    Output: an adapted online prompt used for target-domain prediction.

    1. Initialize the online prompt and target prompt.
    2. Use the frozen model with a hand-crafted prompt to classify the available test images, recording each image’s pseudo-label and confidence.
    3. In offline adaptation, retain the top KK highest-confidence images for each class. In online adaptation, rank the currently available data in each arriving mini-batch and retain the top kk images, with k<Kk<K; recompute the selected adaptation set as new mini-batches arrive.
    4. For each selected image, generate two augmented views using the SimCLR-style augmentation procedure, compute the target-prompt assignments and online-prompt predictions, and evaluate the swapped-prediction and pseudo-label losses.
    5. Update only the online prompt by minimizing the weighted adaptation loss αLswap+βLpseudo\alpha\mathcal{L}_{\mathrm{swap}}+\beta\mathcal{L}_{\mathrm{pseudo}} with SGD.
    6. After every optimization step, update the target prompt using the exponential moving average with decay ϵ\epsilon.
    7. Stop after the chosen adaptation budget, or continue updating as online mini-batches arrive; use the final online prompt for test prediction. The paper reports 50 optimization epochs for its standard offline experiments and does not report an asymptotic computational-complexity bound.
  6. Knowl 6 — Experimental protocol and benchmark coverage

    experimental setup

    SwapPrompt is evaluated using the public CLIP model with a ResNet-50 visual encoder on 14 datasets: Caltech101, DTD, Flowers102, Oxford-Pets, UCF101, StanfordCars, Food101, EuroSAT, SUN397, ImageNet, ImageNet-V2, ImageNet-A, ImageNet-R, and ImageNet-Sketch. Only test images are used for adaptation and evaluation. The compared methods are zero-shot CLIP, UPL, TPT, and supervised CoOp; CoOp receives 4 labeled target-domain examples per class and is treated as an upper-bound reference for test-time adaptation.

    Unless otherwise stated, SwapPrompt uses four randomly initialized learnable context tokens, K=16K=16 selected images per class, target-prompt decay ϵ=0.99\epsilon=0.99, α=β=1\alpha=\beta=1, two SimCLR-style augmented views, SGD for 50 epochs, cosine learning-rate decay, initial learning rate 0.0020.002, and batch size 32. Experiments were run on an RTX 3090 GPU with a 3.5-GHz Intel Core i9-11900K CPU and 64 GB RAM.

  7. Knowl 7 — Test-time accuracy across fourteen datasets

    data/table

    The following comparison reports classification accuracy in percent for zero-shot CLIP, unsupervised prompt learning, entropy-based test-time prompt tuning, supervised few-shot CoOp, SwapPrompt, and online SwapPrompt. The Δ\Delta row is SwapPrompt’s gain over the better of UPL and TPT for each dataset. SwapPrompt improves over both unsupervised baselines on most datasets, approaches or exceeds the supervised CoOp reference on several datasets, and has average accuracy gains of 2.31 percentage points over TPT and 2.17 percentage points over UPL.

    Method Caltech101 DTD Flowers102 Oxford-Pets UCF101 StanfordCars Food101 EuroSAT SUN397 ImageNet ImageNet-V2 ImageNet-A ImageNet-R ImageNet-Sketch
    CoOp 88.76 54.62 83.98 87.44 66.71 61.83 73.79 61.68 64.33 61.23 55.29 23.41 56.96 35.64
    CLIP 85.13 42.16 65.40 83.05 61.15 55.65 74.23 37.60 58.55 58.18 51.36 21.69 55.98 33.33
    UPL 86.37 45.04 67.11 88.53 63.63 58.46 74.38 41.40 61.07 61.19 52.07 23.59 57.09 36.40
    TPT 87.22 42.17 65.42 84.60 61.18 58.49 74.88 43.82 61.46 60.74 54.35 26.24 58.72 35.02
    SwapPrompt 89.90 47.34 70.22 89.14 65.66 59.60 75.08 46.64 63.93 61.80 53.94 24.46 60.88 38.21
    Δ\Delta +2.68 +2.30 +3.11 +0.61 +2.03 +1.11 +0.20 +2.82 +2.47 +0.61 -0.41 -1.78 +2.16 +1.81
    SwapPrompt + Online 89.69 46.40 68.12 88.97 64.52 58.88 75.66 42.45 63.36 61.41 52.93 24.42 60.25 38.13
  8. Knowl 8 — Ablation of augmentation and swapped-prediction objectives

    empirical result

    On ImageNet, Caltech101, DTD, Flowers102, Oxford-Pets, and UCF101, adding two-view augmentation to the pseudo-label baseline improves its mean accuracy from 68.65% to 69.39%, while the complete SwapPrompt objective further improves it to 70.68%. The result indicates that swapped prediction contributes beyond ordinary pseudo-label training.

    Method ImageNet Caltech101 DTD Flowers102 Oxford-Pets UCF101 Average
    UPL 61.19 86.37 45.04 67.11 88.53 63.63 68.65
    UPL+AUG 61.30 87.75 46.04 68.43 87.67 65.15 69.39
    SwapPrompt 61.80 89.90 47.34 70.22 89.14 65.66 70.68

    The paper also compares three cross-view losses on Caltech101, DTD, Flowers102, Oxford-Pets, and UCF101. With online-prompt predictions pi1,pi2p_i^1,p_i^2 and target-prompt assignments qi1,qi2q_i^1,q_i^2, the alternatives are L1=ℓ(pi1,pi2)+ℓ(qi2,qi1)L_1=\ell(p_i^1,p_i^2)+\ell(q_i^2,q_i^1), which removes online-target guidance, and L2=ℓ(pi1,qi1)+ℓ(pi2,qi2)L_2=\ell(p_i^1,q_i^1)+\ell(p_i^2,q_i^2), which removes cross-view swapping. The proposed LswapL_{\mathrm{swap}} obtains the highest mean accuracy, 72.45%, compared with 71.32% for L1L_1 and 71.24% for L2L_2.

    Loss Caltech101 DTD Flowers102 Oxford-Pets UCF101 Average
    L1L_1 87.38 47.22 69.63 87.71 64.68 71.32
    L2L_2 88.45 46.69 69.28 87.30 64.46 71.24
    LswapL_{\mathrm{swap}} 89.90 47.34 70.22 89.14 65.66 72.45
  9. Knowl 9 — Effect of target-prompt decay

    empirical result

    The EMA decay rate controls the trade-off between rapidly updating the target prompt and retaining historical prompt information. Across ImageNet, Caltech101, DTD, Flowers102, Oxford-Pets, and UCF101, intermediate decay rates are consistently better than the extremes ϵ=0\epsilon=0 and ϵ=1\epsilon=1. The best mean accuracy is 70.68% at ϵ=0.99\epsilon=0.99; fixing the target prompt at ϵ=1\epsilon=1 gives 69.33%, so EMA updating improves the mean by 1.35 percentage points.

    ϵ\epsilon ImageNet Caltech101 DTD Flowers102 Oxford-Pets UCF101 Average
    1 60.87 88.49 45.21 68.57 87.01 65.85 69.33
    0.999 61.88 89.71 47.58 68.70 88.74 66.93 70.59
    0.99 61.80 89.90 47.34 70.22 89.14 65.66 70.68
    0.9 61.76 87.99 47.64 70.65 87.47 65.00 70.09
    0 61.29 87.38 47.22 69.63 87.71 64.68 69.65

    At ϵ=0\epsilon=0, the target prompt immediately equals the online prompt, eliminating the historical delay and causing the swapped-prediction setup to suffer gradient collapse when the prompts are effectively identical. At ϵ=1\epsilon=1, the target prompt remains at its initialization and supplies only fixed zero-shot information.

  10. Knowl 10 — Practical effects of data selection, adaptation speed, and prompt configuration

    empirical result

    Confidence-based selection is important because high-confidence pseudo-labels can still be noisy. On Caltech101, DTD, Flowers102, Oxford-Pets, and UCF101, selecting K=16K=16 images per class gives the best mean accuracy, 72.45%, whereas using the entire test set without selection gives 69.95%. The selected-data results are:

    KK Caltech101 DTD Flowers102 Oxford-Pets UCF101 Average
    1 87.51 41.13 65.04 87.83 62.20 68.74
    2 88.72 43.50 64.66 87.42 62.97 69.45
    4 89.49 44.15 66.63 88.98 62.76 70.40
    8 90.14 47.70 66.30 88.43 63.96 71.31
    16 89.90 47.34 70.22 89.14 65.66 72.45
    None 88.03 43.97 68.41 86.68 62.64 69.95

    Adaptation accuracy rises rapidly during the first three epochs and reaches its maximum around epoch 20 before stabilizing. SwapPrompt surpasses the final UPL accuracy after two epochs and the final TPT accuracy after one epoch, allowing a speed–accuracy trade-off when test-time computation is limited. Longer prompts do not improve the mean accuracy in the tested range: context lengths 4, 8, and 16 produce mean accuracies of 72.45%, 72.18%, and 71.83%, respectively. Prompt initialization has little effect; hand-crafted, ImageNet-pretrained, and random initialization produce mean accuracies of 72.33%, 72.36%, and 72.45%, respectively. Varying (α,β)(\alpha,\beta) among (0.6,1.4)(0.6,1.4), (1.0,1.0)(1.0,1.0), and (1.4,0.6)(1.4,0.6) gives mean accuracies of 71.77%, 72.45%, and 72.12%, indicating limited sensitivity. In the online setting, performance is usually slightly lower because early mini-batches contain too little data to learn a suitable prompt; high-confidence pseudo-label noise can also make an intermediate prompt outperform the final prompt on some datasets.

Coverage note — No substantial contributed material was omitted; minor implementation hardware details are included in the experimental protocol, while the paper’s reference and background material is excluded.

References

  1. 1.J Quinonero Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence. Dataset shift in machine learning. The MIT Press, 1:5, 2009.
  2. 2.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In International Conference on Learning Representations, 2019.
  3. 3.Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. Wilds: A benchmark of in-the-wild distribution shifts. In International Conference on Machine Learning, pages 5637–5664. PMLR, 2021.
  4. 4.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In Proceedings of the European conference on computer vision (ECCV), pages 132–149, 2018.
  5. 5.Judy Hoffman, Eric Tzeng, Taesung Park, Jun-Yan Zhu, Phillip Isola, Kate Saenko, Alexei Efros, and Trevor Darrell. Cycada: Cycle-consistent adversarial domain adaptation. In International conference on machine learning, pages 1989–1998. Pmlr, 2018.
  6. 6.Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, and Moritz Hardt. Test-time training with self-supervision for generalization under distribution shifts. In International conference on machine learning, pages 9229–9248. PMLR, 2020.
  7. 7.Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. Tent: Fully test-time adaptation by entropy minimization. In International Conference on Learning Representations, 2021.
  8. 8.Dian Chen, Dequan Wang, Trevor Darrell, and Sayna Ebrahimi. Contrastive test-time adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 295–305, 2022.
  9. 9.Malik Boudiaf, Romain Mueller, Ismail Ben Ayed, and Luca Bertinetto. Parameter-free online test-time adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8344–8353, 2022.
  10. 10.Fatemeh Azimi, Sebastian Palacio, Federico Raue, Jörn Hees, Luca Bertinetto, and Andreas Dengel. Self-supervised test-time adaptation on video data. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3439–3448, 2022.
  11. 11.Chaithanya Kumar Mummadi, Robin Hutmacher, Kilian Rambach, Evgeny Levinkov, Thomas Brox, and Jan Hendrik Metzen. Test-time adaptation to distribution shift by confidence maximization and input transformation. arXiv preprint arXiv:2106.14999, 2021.
  12. 12.Jian Liang, Dapeng Hu, and Jiashi Feng. Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation. In International Conference on Machine Learning, pages 6028–6039. PMLR, 2020.
  13. 13.Yusuke Iwasawa and Yutaka Matsuo. Test-time classifier adjustment module for model-agnostic domain generalization. Advances in Neural Information Processing Systems, 34:2427–2440, 2021.
  14. 14.Jogendra Nath Kundu, Naveen Venkat, R Venkatesh Babu, et al. Universal source-free domain adaptation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4544–4553, 2020.
  15. 15.Rui Li, Qianfen Jiao, Wenming Cao, Hau-San Wong, and Si Wu. Model adaptation: Unsupervised domain adaptation without source data. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9641–9650, 2020.
  16. 16.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  17. 17.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9):2337–2348, 2022.
  18. 18.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16816–16825, 2022.
  19. 19.Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, and Chaowei Xiao. Test-time prompt tuning for zero-shot generalization in vision-language models. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho, editors, Advances in Neural Information Processing Systems, 2022.
  20. 20.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. Advances in neural information processing systems, 33:9912–9924, 2020.
  21. 21.Connor Shorten, Taghi M Khoshgoftaar, and Borko Furht. Text data augmentation for deep learning. Journal of big Data, 8:1–34, 2021.
  22. 22.Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. Clip-adapter: Better vision-language models with feature adapters. arXiv preprint arXiv:2110.04544, 2021.
  23. 23.Renrui Zhang, Rongyao Fang, Wei Zhang, Peng Gao, Kunchang Li, Jifeng Dai, Yu Qiao, and Hongsheng Li. Tip-adapter: Training-free clip-adapter for better vision-language modeling. arXiv preprint arXiv:2111.03930, 2021.
  24. 24.Tony Huang, Jack Chu, and Fangyun Wei. Unsupervised prompt learning for vision-language models. arXiv preprint arXiv:2204.03649, 2022.
  25. 25.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  26. 26.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400. PMLR, 2019.
  27. 27.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15262–15271, 2021.
  28. 28.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8340–8349, 2021.
  29. 29.Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. Advances in Neural Information Processing Systems, 32, 2019.
  30. 30.Li Fei-Fei, Rob Fergus, and Pietro Perona. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In 2004 conference on computer vision and pattern recognition workshop, pages 178–178. IEEE, 2004.
  31. 31.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3606–3613, 2014.
  32. 32.Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, pages 722–729. IEEE, 2008.
  33. 33.Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498–3505. IEEE, 2012.
  34. 34.Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012.
  35. 35.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pages 554–561, 2013.
  36. 36.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VI 13, pages 446–461. Springer, 2014.
  37. 37.Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217–2226, 2019.
  38. 38.Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conference on computer vision and pattern recognition, pages 3485–3492. IEEE, 2010.
  39. 39.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  40. 40.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.

Citation

MLA
MA, X., et al. “SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 65252–64, https://proceedings.neurips.cc/paper_files/paper/2023/file/cdd0640218a27e9e2c0e52e324e25db0-Paper-Conference.pdf.
APA
MA, X., ZHANG, J., Guo, S., & Xu, W. (2023). SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models. Advances in Neural Information Processing Systems, 36, 65252–65264. https://proceedings.neurips.cc/paper_files/paper/2023/file/cdd0640218a27e9e2c0e52e324e25db0-Paper-Conference.pdf
Chicago
MA, X., J. ZHANG, S. Guo, and W. Xu. 2023. “SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models”. Advances in Neural Information Processing Systems 36: 65252–64. https://proceedings.neurips.cc/paper_files/paper/2023/file/cdd0640218a27e9e2c0e52e324e25db0-Paper-Conference.pdf.
Harvard
MA, X. et al. (2023) “SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 65252–65264. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/cdd0640218a27e9e2c0e52e324e25db0-Paper-Conference.pdf.
Vancouver
1. MA X, ZHANG J, Guo S, Xu W (2023) SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 65252–65264

BibTeX

@inproceedings{ma2023swapprompt,
  title = {SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models},
  author = {MA, XIAOSONG and ZHANG, Jie and Guo, Song and Xu, Wenchao},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {65252-65264},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/cdd0640218a27e9e2c0e52e324e25db0-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors