Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness

Sibo WangJie ZhangZheng YuanShiguang Shan

article2024CVPR80 citations

Proposes a fine-tuning method that aligns adversarial features with representations from the original pre-trained model to prevent overfitting and boost CLIP's zero-shot defense performance on unseen datasets.

Listen

Large vision-language models, such as Contrastive Language-Image Pre-training (CLIP), have demonstrated strong generalization across various computer vision tasks. However, these systems remain vulnerable to adversarial attacks—small, deliberate visual perturbations that mislead model predictions without noticeably altering images. While standard defense mechanisms rely on adversarial fine-tuning on downstream datasets, directly adapting large pre-trained models in this manner causes severe overfitting. Consequently, models lose their broad generalization ability and suffer substantial performance drops on clean, unperturbed images.

The article introduces and evaluates Pre-trained Model Guided Adversarial Fine-Tuning (PMG-AFT), a fine-tuning strategy designed to protect zero-shot vision-language models against adversarial manipulation while preserving their baseline generalization and standard accuracy.

To evaluate this approach, the researchers fine-tuned a standard CLIP vision-language model on a single dataset (such as TinyImageNet) and tested its zero-shot performance across 15 diverse, unseen datasets spanning general object recognition, fine-grained classification, scene recognition, domain-specific tasks, and medical imagery. The PMG-AFT framework establishes an auxiliary generalization branch that aligns the target model's output representations of adversarial images with those of the original, frozen pre-trained model using relative entropy distance constraints, alongside a clean-image consistency regularizer.

The experimental findings show that PMG-AFT significantly outperforms existing defense strategies. Under multi-step adversarial attacks (PGD-10), PMG-AFT achieved an average robust accuracy of 31.95%, reflecting a 4.99% improvement over the previous state-of-the-art defense (FT-TeCoA) and a 14.57% increase over unmodified CLIP. Crucially, PMG-AFT avoided the severe performance penalties typical of adversarial fine-tuning, achieving a clean-image accuracy of 55.71%—an 8.72% improvement over FT-TeCoA. Under strict benchmark testing using AutoAttack, the model maintained superior robustness, achieving 17.91% average accuracy compared to 10.33% for FT-TeCoA and 1.74% for unmodified CLIP.

These results demonstrate that organizations can successfully harden foundation vision models against security exploits without compromising their core commercial utility across zero-shot deployments. By transferring feature constraints directly from the pre-trained model during parameter updates, PMG-AFT effectively resolves the historical trade-off between defensive robustness and clean data accuracy, reducing the operational risk of deploying vision-language systems into safety-critical environments.

Organizations deploying vision-language foundation models in security-sensitive zero-shot environments should adopt feature-guided fine-tuning frameworks like PMG-AFT instead of conventional adversarial training. Decision-makers should account for computational trade-offs, as the dual-branch framework adds approximately 298 seconds of training time per epoch compared to existing baselines. Practitioners adopting the method should implement distance matching at the output prediction layer rather than intermediate feature layers, which produced optimal empirical stability in ablation tests.

While the empirical results demonstrate strong defense gains across standard benchmark distributions, the evaluation is primarily focused on image encoder perturbations within the CLIP architecture. Confidence in the reported zero-shot performance is high across standard visual categories, though practitioners should conduct task-specific pilots when deploying to heavily specialized domains or facing multimodal adversarial attacks.

Cover for Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness

Abstract

Large-scale pre-trained vision-language models like CLIP have demonstrated impressive performance across various tasks, and exhibit remarkable zero-shot generalization capability, while they are also vulnerable to imperceptible adversarial examples. Existing works typically employ adversarial training (fine-tuning) as a defense method against adversarial examples. However, direct application to the CLIP model may result in overfitting, compromising the model's capacity for generalization. In this paper, we propose Pre-trained Model Guided Adversarial Fine-Tuning (PMG-AFT) method, which leverages supervision from the original pre-trained model by carefully designing an auxiliary branch, to enhance the model's zero-shot adversarial robustness. Specifically, PMG-AFT minimizes the distance between the features of adversarial examples in the target model and those in the pre-trained model, aiming to preserve the generalization features already captured by the pre-trained model. Extensive Experiments on 15 zero-shot datasets demonstrate that PMG-AFT significantly outperforms the state-of-the-art method, improving the top-1 robust accuracy by an average of 4.99%. Furthermore, our approach consistently improves clean accuracy by an average of 8.72%. Our code is available at here.1

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Preliminaries and Problem Setup
  • 3.2. Adversarial Fine-Tuning is Prone to Overfitting
  • 3.3. PMG-AFT
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Main Result
  • 4.3. Performance against AutoAttack
  • 4.4. Performance against Different Attack Strength
  • 4.5. Trade-off between Robust and Clean Accuracy
  • 4.6. Ablation Study
  • 5. Conclusion
  • 6. Acknowledgement
  • References

Knowls

  1. Knowl 1 — Pre-trained Model Guided Adversarial Fine-Tuning

    model/method

    Pre-trained Model Guided Adversarial Fine-Tuning (PMG-AFT) adapts CLIP for zero-shot adversarial robustness by constraining a trainable target image encoder with information from the original pre-trained CLIP. The target image encoder FθF_{\theta} is initialized from the original image encoder ForiF_{\mathrm{ori}}; the original image encoder and the original text encoder ToriT_{\mathrm{ori}} remain frozen, and only θ\theta is updated.

    For cc categories, the frozen text encoder converts one prompt per category into a matrix T∈Rc×dT\in\mathbb{R}^{c\times d}, where dd is the embedding dimension. For each minibatch, PMG-AFT generates adversarial images with the current target encoder and the frozen text embeddings. It then updates the target encoder using three signals: classification supervision on adversarial images, agreement between the target and original CLIP on adversarial images, and agreement between the target encoder's adversarial and clean predictions.

    The method alternates adversarial-example generation with target-encoder parameter updates. The generalization branch uses the original CLIP as a fixed teacher in output-probability space, rather than constraining the target model directly in parameter space.

  2. Knowl 2 — PMG-AFT training objective

    equation

    Let XX be a minibatch of NN clean images, Y∈{0,1}N×cY\in\{0,1\}^{N\times c} their one-hot labels, XaX_a the corresponding adversarial images, and T∈Rc×dT\in\mathbb{R}^{c\times d} the frozen matrix of category text embeddings. Define target-model image embeddings I=Fθ(X)I=F_{\theta}(X) and Ia=Fθ(Xa)I_a=F_{\theta}(X_a), and original-model adversarial embeddings Iori,a=Fori(Xa)I_{\mathrm{ori},a}=F_{\mathrm{ori}}(X_a). The three row-wise class-probability matrices are

    Pa=softmax⁡(IaT⊤),Pori,a=softmax⁡(Iori,aT⊤),Pc=softmax⁡(IT⊤).P_a=\operatorname{softmax}(I_aT^{\top}),\qquad P_{\mathrm{ori},a}=\operatorname{softmax}(I_{\mathrm{ori},a}T^{\top}),\qquad P_c=\operatorname{softmax}(IT^{\top}).

    For a probability vector PP, let DKL(P∥Q)=∑r=1cPrlog⁡(Pr/Qr)D_{\mathrm{KL}}(P\|Q)=\sum_{r=1}^{c}P_r\log(P_r/Q_r). PMG-AFT minimizes

    L=Lrobust+αLgeneral+βLclean,L=L_{\mathrm{robust}}+\alpha L_{\mathrm{general}}+\beta L_{\mathrm{clean}},

    where α,β≥0\alpha,\beta\geq 0 are loss weights and

    Lrobust=−1N∑i=1N∑r=1cYirlog⁡(Pa,ir),L_{\mathrm{robust}}=-\frac{1}{N}\sum_{i=1}^{N}\sum_{r=1}^{c}Y_{ir}\log(P_{a,ir}), Lgeneral=1N∑i=1NDKL(Pa,i∥Pori,a,i),Lclean=1N∑i=1NDKL(Pa,i∥Pc,i).L_{\mathrm{general}}=\frac{1}{N}\sum_{i=1}^{N}D_{\mathrm{KL}}(P_{a,i}\|P_{\mathrm{ori},a,i}),\qquad L_{\mathrm{clean}}=\frac{1}{N}\sum_{i=1}^{N}D_{\mathrm{KL}}(P_{a,i}\|P_{c,i}).

    Here Pa,iP_{a,i}, Pori,a,iP_{\mathrm{ori},a,i}, and Pc,iP_{c,i} are the class-probability vectors for sample ii. The first term makes adversarial examples predict their ground-truth classes; the second transfers the original CLIP's predictions on adversarial images; and the third encourages adversarial predictions to remain close to the target model's predictions on clean images. The reported experiments use α=1\alpha=1 and β=1\beta=1.

  3. Knowl 3 — Zero-shot adversarial robustness evaluation

    definition

    The paper defines zero-shot adversarial robustness by fine-tuning CLIP on a target dataset and then testing adversarial examples from other datasets without further adaptation. Specifically, a target model is adversarially fine-tuned on a labeled dataset such as TinyImageNet. At test time, white-box adversarial examples are generated for each evaluation dataset using the evaluated model, and robust accuracy is the percentage of those adversarial images classified correctly.

    Clean accuracy is measured on the corresponding unperturbed images. The TinyImageNet evaluation is not strictly zero-shot because TinyImageNet is used for fine-tuning; the other datasets measure transfer of adversarial robustness to unseen data distributions and categories.

  4. Knowl 4 — Main zero-shot robustness and clean-accuracy results

    data/table

    The main comparison reported on page 7 fine-tunes CLIP on TinyImageNet and evaluates all methods on 16 datasets with PGD-10 at ϵ=1/255\epsilon=1/255. The table reports mean accuracy over the 16 datasets and the reported time in seconds. The primary fine-tuning comparison shows that PMG-AFT improves robust accuracy while avoiding most of the clean-accuracy degradation caused by adversarial fine-tuning.

    Method Average robust accuracy (%) Average clean accuracy (%) Time (s)
    CLIP 17.38 59.98 0
    FT-Standard 13.79 54.76 234
    FT-TeCoA 26.96 46.99 551
    PMG-AFT 31.95 55.71 849
    VP-TeCoA 18.27 28.50 364
    VPT-PMG-AFT 22.74 36.03 661

    Relative to the original CLIP, PMG-AFT raises average PGD-10 robust accuracy by 14.57 percentage points. Relative to the previous adversarial fine-tuning method FT-TeCoA, it raises robust accuracy by 4.99 points and clean accuracy by 8.72 points. PMG-AFT improves robust accuracy over FT-TeCoA on most datasets, although its TinyImageNet robust accuracy is lower because TinyImageNet is the fine-tuning dataset. Its clean accuracy is close to the original CLIP and exceeds FT-Standard by 0.95 points.

  5. Knowl 5 — Experimental protocol and datasets

    experimental setup

    The experiments use CLIP ViT-B/32 as the backbone. The image encoder is adversarially fine-tuned for 10 epochs with SGD and learning rate 5×10−55\times10^{-5} while updating all image-encoder parameters; the text encoder is frozen. Visual-prompt baselines use a learning rate of 4040, an image-layer prompt with the same dimensions as the input image, and a token-layer prompt of dimension 100. Training uses two NVIDIA GeForce RTX 3090 GPUs.

    The model is fine-tuned primarily on TinyImageNet and evaluated on TinyImageNet plus 15 additional datasets: CIFAR10, CIFAR100, STL10, ImageNet, Caltech101, Caltech256, OxfordPets, Flowers102, FGVCAircraft, StanfordCars, SUN397, Food101, EuroSAT, DTD, and PCAM. All images are resized or preprocessed to 3×224×2243\times224\times224 inputs.

    Adversarial training and evaluation use ℓ∞\ell_{\infty} PGD-2 and PGD-10 with perturbation bounds ϵ∈{1/255,2/255,4/255}\epsilon\in\{1/255,2/255,4/255\}. Unless otherwise stated, the principal comparison uses PGD-10 with ϵ=1/255\epsilon=1/255 for both training and testing. The PMG-AFT loss weights are α=1\alpha=1 and β=1\beta=1.

  6. Knowl 6 — Adversarial fine-tuning exhibits generalization overfitting

    empirical result

    The paper's diagnostic experiments show that adversarial fine-tuning can move CLIP too far from the general-purpose solution learned during pre-training. Standard fine-tuning on clean TinyImageNet improves performance on TinyImageNet but decreases accuracy on other zero-shot datasets. FT-TeCoA improves adversarial accuracy across many datasets, but its clean accuracy decreases substantially, indicating that robustness learned from the target dataset does not transfer cleanly to unseen data.

    The parameter-distance analysis shown on page 5 tracks relative L2L_2 distance from the original CLIP during training. Both clean fine-tuning and adversarial fine-tuning depart from the original parameters, but adversarial fine-tuning departs farther. The paper uses this observation to motivate feature- and output-space guidance from the original model rather than relying only on parameter-space regularization.

  7. Knowl 7 — Robustness under AutoAttack and stronger perturbations

    empirical result

    PMG-AFT remains stronger than the original CLIP and FT-TeCoA under attacks beyond the PGD-10 evaluation. On the 16 datasets with standard AutoAttack at ϵ=1/255\epsilon=1/255, the mean robust accuracies reported on page 8 are 1.74% for CLIP, 10.33% for FT-TeCoA, and 17.91% for PMG-AFT. Thus, although PMG-AFT also loses accuracy when the attack is strengthened from PGD-10 to AutoAttack, it retains the best mean robustness.

    When the training and testing perturbation bounds are matched and increased from 1/2551/255 to 2/2552/255 and 4/2554/255, robust accuracy decreases for both PMG-AFT and FT-TeCoA, but PMG-AFT remains consistently better on the sampled datasets. The clean accuracy is reported to be nearly unaffected by this increase in perturbation bound. In the clean-versus-robust comparison, PMG-AFT achieves the best joint average of robust accuracy and clean accuracy among the compared fine-tuning methods.

  8. Knowl 8 — Loss-term ablation

    data/table

    The loss ablation reported on page 8 isolates the contributions of the original-model guidance and clean-output regularization. All values are averages over the evaluation datasets and are percentages.

    Method Average clean accuracy (%) Average robust accuracy (%)
    FT-TeCoA (α=0,β=0)(\alpha=0,\beta=0) 46.99 26.96
    +Lgeneral+L_{\mathrm{general}} (α=1,β=0)(\alpha=1,\beta=0) 58.28 28.44
    +Lgeneral+Lclean+L_{\mathrm{general}}+L_{\mathrm{clean}} (α=1,β=1)(\alpha=1,\beta=1) 55.72 32.04

    Adding the generalization-information loss to FT-TeCoA raises mean robust accuracy by 1.48 points and mean clean accuracy by 11.29 points. Adding the clean-output regularizer then raises mean robust accuracy from 28.44% to 32.04%. The reported hyperparameter sweep finds the best overall setting at α=1\alpha=1 and β=1\beta=1.

  9. Knowl 9 — Output-space KL guidance is the most effective variant

    data/table

    PMG-AFT was evaluated using either the model output or the penultimate feature as the guided representation, and using KL divergence, L2L_2 distance, or cosine distance. The results reported on page 8 show that comparing output probability distributions with KL divergence is best for both clean and robust accuracy.

    Guided representation and distance Average clean accuracy (%) Average robust accuracy (%)
    Output + KL 55.71 31.95
    Output + L2L_2 48.21 27.23
    Feature + L2L_2 38.32 22.90
    Feature + cosine 52.98 24.53

    Thus, the reported PMG-AFT configuration distills the original CLIP's class-probability output on adversarial images, rather than matching penultimate features or using a distance other than KL divergence.

  10. Knowl 10 — Transfer to CIFAR100 and computational overhead

    empirical result

    When CIFAR100 is used as the fine-tuning dataset instead of TinyImageNet and the same evaluation protocol is applied to the 16 datasets, PMG-AFT obtains an average robust accuracy of 26.06%, which is 4.71 percentage points higher than FT-TeCoA.

    The additional original-model branch increases computational cost. Relative to FT-TeCoA, the paper reports an extra 298 seconds per training epoch for PMG-AFT. This overhead is the principal stated practical limitation of adding the generalization-information branch, although the method produces higher robust and clean accuracy in the reported experiments.

Coverage note — The complete per-dataset cells of the 16-dataset PGD and AutoAttack tables, and supplementary per-dataset attack-strength results, were omitted because the aggregate comparisons and ablations capture their contribution without adding separate load-bearing findings.

References

  1. 1.Anish Athalye, Nicholas Carlini, and David Wagner. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International conference on machine learning, pages 274–283. PMLR, 2018.
  2. 2.Randall Balestriero, Mark Ibrahim, Vlad Sobal, Ari Morcos, Shashank Shekhar, Tom Goldstein, Florian Bordes, Adrien Bardes, Gregoire Mialon, Yuandong Tian, et al. A cookbook of self-supervised learning. arXiv preprint arXiv:2304.12210, 2023.
  3. 3.Babak Ehteshami Bejnordi, Mitko Veta, Paul Johannes Van Diest, Bram Van Ginneken, Nico Karssemeijer, Geert Litjens, Jeroen AWM Van Der Laak, Meyke Hermsen, Quirine F Manson, Maschenka Balkenhol, et al. Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer. Jama, 318(22):2199–2210, 2017.
  4. 4.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VI 13, pages 446–461. Springer, 2014.
  5. 5.Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In 2017 ieee symposium on security and privacy (sp), pages 39–57. Ieee, 2017.
  6. 6.Mathilde Caron, Hugo Touvron, Ishan Misra, Herve J ´ egou, ´ Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021.
  7. 7.Jun Chen, Han Guo, Kai Yi, Boyang Li, and Mohamed Elhoseiny. Visualgpt: Data-efficient adaptation of pretrained language models for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18030–18040, 2022.
  8. 8.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  9. 9.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3606–3613, 2014.
  10. 10.Adam Coates, Andrew Y Ng, and Honglak Lee. An analysis of single-layer networks in unsupervised feature learning. AISTATS, 15:215–223, 2011.
  11. 11.Francesco Croce and Matthias Hein. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In International conference on machine learning, pages 2206–2216. PMLR, 2020.
  12. 12.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  14. 14.Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2002.06305, 2020.
  15. 15.Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial attacks with momentum. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 9185–9193, 2018.
  16. 16.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  17. 17.Li Fei-Fei, Robert Fergus, and Pietro Perona. One-shot learning of object categories. IEEE transactions on pattern analysis and machine intelligence, 28(4):594–611, 2006.
  18. 18.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572, 2014.
  19. 19.Gregory Griffin, Alex Holub, and Pietro Perona. Caltech-256 object category dataset. 2007.
  20. 20.Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217–2226, 2019.
  21. 21.Nathan Inkawhich, Gwendolyn McDonald, and Ryan Luley. Adversarial attacks on foundational vision models. arXiv preprint arXiv:2308.14597, 2023.
  22. 22.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR, 2021.
  23. 23.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pages 554–561, 2013.
  24. 24.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  25. 25.Cheolhyoung Lee, Kyunghyun Cho, and Wanmo Kang. Mixout: Effective regularization to finetune large-scale pretrained language models. arXiv preprint arXiv:1909.11299, 2019.
  26. 26.Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven CH Hoi. Align and prompt: Video-and-language pre-training with entity prompts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4953–4963, 2022.
  27. 27.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, pages 12888–12900. PMLR, 2022.
  28. 28.Xiao Li, Wei Zhang, Yining Liu, Zhanhao Hu, Bo Zhang, and Xiaolin Hu. Anchor-based adversarially robust zero-shot learning driven by language. arXiv preprint arXiv:2301.13096, 2023.
  29. 29.Dong Lu, Zhiqiang Wang, Teng Wang, Weili Guan, Hongchang Gao, and Feng Zheng. Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 102–111, 2023.
  30. 30.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083, 2017.
  31. 31.Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151, 2013.
  32. 32.Chengzhi Mao, Scott Geng, Junfeng Yang, Xin Wang, and Carl Vondrick. Understanding zero-shot adversarial robustness for large-scale models. In The Eleventh International Conference on Learning Representations, 2023.
  33. 33.Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In 2008 Sixth Indian conference on computer vision, graphics & image processing, pages 722–729. IEEE, 2008.
  34. 34.Maxime Oquab, Timothee Darcet, Th ´ eo Moutakanni, Huy ´ Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023.
  35. 35.Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498–3505. IEEE, 2012.
  36. 36.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  37. 37.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021.
  38. 38.Mr D Murahari Reddy, Mr Sk Masthan Basha, Mr M Chinnaiahgari Hari, and Mr N Penchalaiah. Dall-e: Creating images from text. UGC Care Group I Journal, 8(14):71–75, 2021.
  39. 39.Leslie Rice, Eric Wong, and Zico Kolter. Overfitting in adversarially robust deep learning. In International Conference on Machine Learning, pages 8093–8104. PMLR, 2020.
  40. 40.Ali Shafahi, Mahyar Najibi, Mohammad Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein. Adversarial training for free! Advances in Neural Information Processing Systems, 32, 2019.
  41. 41.Mannat Singh, Laura Gustafson, Aaron Adcock, Vinicius de Freitas Reis, Bugra Gedik, Raj Prateek Kosaraju, Dhruv Mahajan, Ross Girshick, Piotr Dollar, and Laurens Van Der Maaten. Revisiting weakly supervised pre-training of visual perception models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 804–814, 2022.
  42. 42.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530, 2019.
  43. 43.Yifei Wang, Liangchen Li, Jiansheng Yang, Zhouchen Lin, and Yisen Wang. Balance, imbalance, and rebalance: Understanding robust overfitting from a minimax game perspective. Advances in neural information processing systems, 36, 2023.
  44. 44.Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, et al. Robust fine-tuning of zero-shot models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7959–7971, 2022.
  45. 45.Dongxian Wu, Shu-Tao Xia, and Yisen Wang. Adversarial weight perturbation helps robust generalization. Advances in Neural Information Processing Systems, 33:2958–2969, 2020.
  46. 46.Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conference on computer vision and pattern recognition, pages 3485–3492. IEEE, 2010.
  47. 47.Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan. Theoretically principled trade-off between robustness and accuracy. In International conference on machine learning, pages 7472–7482. PMLR, 2019.
  48. 48.Jiaming Zhang, Qi Yi, and Jitao Sang. Towards adversarial attack on vision-language pre-training models. In Proceedings of the 30th ACM International Conference on Multimedia, pages 5005–5013, 2022.
  49. 49.Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi. Revisiting few-sample bert fine-tuning. arXiv preprint arXiv:2006.05987, 2020.
  50. 50.Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Cheung, and Min Lin. On evaluating adversarial robustness of large vision-language models. arXiv preprint arXiv:2305.16934, 2023.

Citation

MLA
Wang, S., et al. “Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness”. arXiv, 2024, http://arxiv.org/abs/2401.04350v3.
APA
Wang, S., Zhang, J., Yuan, Z., & Shan, S. (2024). Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness. arXiv. http://arxiv.org/abs/2401.04350v3
Chicago
Wang, S., J. Zhang, Z. Yuan, and S. Shan. 2024. “Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness”. arXiv. http://arxiv.org/abs/2401.04350v3.
Harvard
Wang, S. et al. (2024) “Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.04350v3.
Vancouver
1. Wang S, Zhang J, Yuan Z, Shan S (2024) Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness. arXiv

BibTeX

@article{wang2024pre,
  title = {Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness},
  author = {Wang, Sibo and Zhang, Jie and Yuan, Zheng and Shan, Shiguang},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.04350v3},
  eprint = {2401.04350}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE