Robust fine-tuning of zero-shot models
Mitchell WortsmanGabriel IlharcoJong Wook KimMike LiSimon KornblithRebecca RoelofsRaphael Gontijo LopesHannaneh HajishirziAli FarhadiHongseok Namkoong
Introduces WiSE-FT, a simple technique that linearly interpolates the weights of zero-shot and fine-tuned models to significantly improve out-of-distribution generalization without sacrificing target accuracy or adding computational overhead.
Modern vision systems increasingly rely on large pre-trained foundation models such as CLIP and ALIGN to perform zero-shot classification across diverse tasks. While these zero-shot models demonstrate remarkable robustness to natural distribution shifts, tailoring them to specific downstream applications via standard fine-tuning creates a severe trade-off. Fine-tuning markedly improves accuracy on the target dataset but consistently degrades robustness on shifted or out-of-distribution real-world data, creating operational risks when deployment environments encounter natural variations in style, geography, or capture conditions.
The article evaluates and demonstrates a simple technique called Weight-Space Ensembles for Fine-Tuning (WiSE-FT) to overcome this limitation. The primary objective is to develop a fine-tuning strategy that enhances robustness under distribution shift while preserving or improving high target-distribution accuracy without incurring extra computational costs.
To address this challenge, the authors propose a two-step approach: first, standard fine-tuning is conducted on the downstream target dataset; second, the original pre-trained zero-shot model weights and the fine-tuned model weights are blended together via linear interpolation using a mixing coefficient. The authors conduct extensive empirical evaluations across multiple model architectures (including CLIP, BASIC, ALIGN, and JFT-pre-trained Vision Transformers) evaluated on standard target benchmarks like ImageNet alongside eleven natural distribution shifts involving geographic shifts, video perturbations, sketches, and diverse image renditions.
The article establishes several key findings. First, WiSE-FT delivers substantial accuracy gains under distribution shift: on ImageNet and five derived shifts, it improves shifted accuracy by 4 to 6 percentage points over prior fine-tuning approaches while boosting target ImageNet accuracy by 1.6 percentage points. Second, across six additional real-world distribution shifts—including satellite imagery, wildlife monitoring, and video datasets—WiSE-FT provides robustness gains ranging from 2 to 23 percentage points relative to standard fine-tuning. Third, even on standard transfer learning benchmarks without explicit distribution shifts, WiSE-FT outperforms standard fine-tuning, reducing relative classification error rates by 4% to 49% across seven datasets. Finally, standard fine-tuning is highly brittle to hyperparameter choices like learning rate and epoch count, whereas WiSE-FT reliably eliminates the trade-off between target accuracy and robustness across varied settings.
These findings indicate that teams deploying machine learning systems do not need to choose between specialized in-distribution performance and real-world robustness. Because weight-space ensembling merges parameters into a single neural network, all performance gains are achieved with zero additional latency, memory footprint, or inference computational overhead compared to a standard deployed model. This dramatically improves reliability in safety-critical and variable environments at no extra operational cost.
For practitioners fine-tuning pre-trained zero-shot vision models, the article recommends adopting WiSE-FT as a standard fine-tuning practice. Setting the mixing coefficient to 0.5 provides near-optimal performance across diverse applications when no specific validation data for distribution shift is available. Practitioners should avoid costly and fragile hyperparameter searches aimed at preserving robustness, and instead tune the linear mixing coefficient directly using the fine-tuned and base weights.
The conclusions are supported by extensive empirical validation across multiple model scales and benchmark datasets. However, the study focuses exclusively on image classification tasks, leaving applications in object detection, segmentation, and natural language processing to future research. Decision-makers can have high confidence in applying this technique to visual recognition tasks, while bearing in mind that downstream deployments still inherit the broader behavioral biases present in the underlying pre-trained foundation models.
- Paper: Averaging Weights Leads to Wider Optima and Better Generalization, Pavel Izmailov et al. (2018). Its Stochastic Weight Averaging establishes how averaging neural-network parameters can improve generalization without the inference costs of a prediction ensemble—the key intuition behind WiSE-FT.
- Paper: Big Transfer (BiT): General Visual Representation Learning, Alexander Kolesnikov et al. (2019). Big Transfer lays out the pretraining-and-fine-tuning paradigm and large-scale visual representations that make the source’s adaptation of pretrained models intelligible.
- Paper: Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP), Alex Fang et al. (2022). Its controlled analysis of what gives CLIP robustness supplies essential context for why preserving a pretrained model’s broad capabilities during fine-tuning matters.
- Paper: WILDS: A Benchmark of in-the-Wild Distribution Shifts, Pang Wei Koh et al. (2020). WILDS establishes realistic distribution-shift evaluation across deployed settings, helping frame the source’s robustness tests beyond ordinary in-distribution accuracy.
- Paper: Trainable Projected Gradient Method for Robust Fine-Tuning, Junjiao Tian et al. (2023). TPGM carries the robust-fine-tuning challenge forward by replacing WiSE-FT’s global weight blend with learned, layer-specific constraints on how far fine-tuned weights move from pretraining.
