Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness
Sibo WangJie ZhangZheng YuanShiguang Shan
Proposes a fine-tuning method that aligns adversarial features with representations from the original pre-trained model to prevent overfitting and boost CLIP's zero-shot defense performance on unseen datasets.
Large vision-language models, such as Contrastive Language-Image Pre-training (CLIP), have demonstrated strong generalization across various computer vision tasks. However, these systems remain vulnerable to adversarial attacks—small, deliberate visual perturbations that mislead model predictions without noticeably altering images. While standard defense mechanisms rely on adversarial fine-tuning on downstream datasets, directly adapting large pre-trained models in this manner causes severe overfitting. Consequently, models lose their broad generalization ability and suffer substantial performance drops on clean, unperturbed images.
The article introduces and evaluates Pre-trained Model Guided Adversarial Fine-Tuning (PMG-AFT), a fine-tuning strategy designed to protect zero-shot vision-language models against adversarial manipulation while preserving their baseline generalization and standard accuracy.
To evaluate this approach, the researchers fine-tuned a standard CLIP vision-language model on a single dataset (such as TinyImageNet) and tested its zero-shot performance across 15 diverse, unseen datasets spanning general object recognition, fine-grained classification, scene recognition, domain-specific tasks, and medical imagery. The PMG-AFT framework establishes an auxiliary generalization branch that aligns the target model's output representations of adversarial images with those of the original, frozen pre-trained model using relative entropy distance constraints, alongside a clean-image consistency regularizer.
The experimental findings show that PMG-AFT significantly outperforms existing defense strategies. Under multi-step adversarial attacks (PGD-10), PMG-AFT achieved an average robust accuracy of 31.95%, reflecting a 4.99% improvement over the previous state-of-the-art defense (FT-TeCoA) and a 14.57% increase over unmodified CLIP. Crucially, PMG-AFT avoided the severe performance penalties typical of adversarial fine-tuning, achieving a clean-image accuracy of 55.71%—an 8.72% improvement over FT-TeCoA. Under strict benchmark testing using AutoAttack, the model maintained superior robustness, achieving 17.91% average accuracy compared to 10.33% for FT-TeCoA and 1.74% for unmodified CLIP.
These results demonstrate that organizations can successfully harden foundation vision models against security exploits without compromising their core commercial utility across zero-shot deployments. By transferring feature constraints directly from the pre-trained model during parameter updates, PMG-AFT effectively resolves the historical trade-off between defensive robustness and clean data accuracy, reducing the operational risk of deploying vision-language systems into safety-critical environments.
Organizations deploying vision-language foundation models in security-sensitive zero-shot environments should adopt feature-guided fine-tuning frameworks like PMG-AFT instead of conventional adversarial training. Decision-makers should account for computational trade-offs, as the dual-branch framework adds approximately 298 seconds of training time per epoch compared to existing baselines. Practitioners adopting the method should implement distance matching at the output prediction layer rather than intermediate feature layers, which produced optimal empirical stability in ablation tests.
While the empirical results demonstrate strong defense gains across standard benchmark distributions, the evaluation is primarily focused on image encoder perturbations within the CLIP architecture. Confidence in the reported zero-shot performance is high across standard visual categories, though practitioners should conduct task-specific pilots when deploying to heavily specialized domains or facing multimodal adversarial attacks.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. Introduces the CLIP vision-language architecture and zero-shot transfer paradigm that forms the core foundation and pre-trained baseline evaluated in the source paper.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, Aleksander Madry et al. (2017). Establishes the standard minimax robust optimization and adversarial training framework that the source adapts to prevent overfitting during vision-language fine-tuning.
- Paper: Trainable Projected Gradient Method for Robust Fine-Tuning, Junjiao Tian et al. (2023). Formulates robust fine-tuning as a constrained optimization problem to preserve pre-trained generalization capabilities under distribution shifts, motivating the source's auxiliary guidance approach.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). Provides fundamental theoretical insights into the trade-offs between robust and non-robust features in deep neural representations during fine-tuning.
- Paper: Adversarial Machine Learning at Scale, Alexey Kurakin et al. (2016). Pioneered large-scale adversarial training on vision datasets, demonstrating the challenges of scaling robust defenses to large architectures like pre-trained models.
- Paper: BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning, Siyuan Liang et al. (2024). Explores backdoor vulnerabilities resistant to fine-tuning defenses in multimodal contrastive learning, presenting a complementary adversarial threat model to the source's defense framework.
- Paper: Visual Adversarial Examples Jailbreak Aligned Large Language Models, Xiangyu Qi et al. (2024). Extends the study of visual adversarial vulnerabilities from classification encoders to multimodal large language models and downstream safety jailbreaks.
