Robust fine-tuning of zero-shot models

Mitchell WortsmanGabriel IlharcoJong Wook KimMike LiSimon KornblithRebecca RoelofsRaphael Gontijo LopesHannaneh HajishirziAli FarhadiHongseok Namkoong

article2022CVPR1,041 citations

Introduces WiSE-FT, a simple technique that linearly interpolates the weights of zero-shot and fine-tuned models to significantly improve out-of-distribution generalization without sacrificing target accuracy or adding computational overhead.

Listen

Modern vision systems increasingly rely on large pre-trained foundation models such as CLIP and ALIGN to perform zero-shot classification across diverse tasks. While these zero-shot models demonstrate remarkable robustness to natural distribution shifts, tailoring them to specific downstream applications via standard fine-tuning creates a severe trade-off. Fine-tuning markedly improves accuracy on the target dataset but consistently degrades robustness on shifted or out-of-distribution real-world data, creating operational risks when deployment environments encounter natural variations in style, geography, or capture conditions.

The article evaluates and demonstrates a simple technique called Weight-Space Ensembles for Fine-Tuning (WiSE-FT) to overcome this limitation. The primary objective is to develop a fine-tuning strategy that enhances robustness under distribution shift while preserving or improving high target-distribution accuracy without incurring extra computational costs.

To address this challenge, the authors propose a two-step approach: first, standard fine-tuning is conducted on the downstream target dataset; second, the original pre-trained zero-shot model weights and the fine-tuned model weights are blended together via linear interpolation using a mixing coefficient. The authors conduct extensive empirical evaluations across multiple model architectures (including CLIP, BASIC, ALIGN, and JFT-pre-trained Vision Transformers) evaluated on standard target benchmarks like ImageNet alongside eleven natural distribution shifts involving geographic shifts, video perturbations, sketches, and diverse image renditions.

The article establishes several key findings. First, WiSE-FT delivers substantial accuracy gains under distribution shift: on ImageNet and five derived shifts, it improves shifted accuracy by 4 to 6 percentage points over prior fine-tuning approaches while boosting target ImageNet accuracy by 1.6 percentage points. Second, across six additional real-world distribution shifts—including satellite imagery, wildlife monitoring, and video datasets—WiSE-FT provides robustness gains ranging from 2 to 23 percentage points relative to standard fine-tuning. Third, even on standard transfer learning benchmarks without explicit distribution shifts, WiSE-FT outperforms standard fine-tuning, reducing relative classification error rates by 4% to 49% across seven datasets. Finally, standard fine-tuning is highly brittle to hyperparameter choices like learning rate and epoch count, whereas WiSE-FT reliably eliminates the trade-off between target accuracy and robustness across varied settings.

These findings indicate that teams deploying machine learning systems do not need to choose between specialized in-distribution performance and real-world robustness. Because weight-space ensembling merges parameters into a single neural network, all performance gains are achieved with zero additional latency, memory footprint, or inference computational overhead compared to a standard deployed model. This dramatically improves reliability in safety-critical and variable environments at no extra operational cost.

For practitioners fine-tuning pre-trained zero-shot vision models, the article recommends adopting WiSE-FT as a standard fine-tuning practice. Setting the mixing coefficient to 0.5 provides near-optimal performance across diverse applications when no specific validation data for distribution shift is available. Practitioners should avoid costly and fragile hyperparameter searches aimed at preserving robustness, and instead tune the linear mixing coefficient directly using the fine-tuned and base weights.

The conclusions are supported by extensive empirical validation across multiple model scales and benchmark datasets. However, the study focuses exclusively on image classification tasks, leaving applications in object detection, segmentation, and natural language processing to future research. Decision-makers can have high confidence in applying this technique to visual recognition tasks, while bearing in mind that downstream deployments still inherit the broader behavioral biases present in the underlying pre-trained foundation models.

Cover for Robust fine-tuning of zero-shot models

Abstract

Large pre-trained models such as CLIP or ALIGN offer consistent accuracy across a range of data distributions when performing zero-shot inference (i.e., without fine-tuning on a specific dataset). Although existing fine-tuning methods substantially improve accuracy on a given target distribution, they often reduce robustness to distribution shifts. We address this tension by introducing a simple and effective method for improving robustness while fine-tuning: ensembling the weights of the zero-shot and fine-tuned models (WiSE-FT). Compared to standard fine-tuning, WiSE-FT provides large accuracy improvements under distribution shift, while preserving high accuracy on the target distribution. On ImageNet and five derived distribution shifts, WiSE-FT improves accuracy under distribution shift by 4 to 6 percentage points (pp) over prior work while increasing ImageNet accuracy by 1.6 pp. WiSE-FT achieves similarly large robustness gains (2 to 23 pp) on a diverse set of six further distribution shifts, and accuracy gains of 0.8 to 3.3 pp compared to standard fine-tuning on commonly used transfer learning datasets. These improvements come at no additional computational cost during fine-tuning or inference.

Table of Contents

  • 1 Introduction
  • 2 Background and experimental setup
  • 3 Weight-space ensembles for fine-tuning
  • 4 Results
  • 5 Discussion
  • 5.1 Zero-shot and fine-tuned models are complementary
  • 5.2 An error landscape perspective
  • 6 Related work
  • 7 Limitations, impact, and conclusion
  • References
  • A Pseudocode for WiSE-FT
  • B Mixing coefficient
  • C Additional experiments
  • C.1 Breakdown of CLIP experiments on ImageNet
  • C.2 Robustness on additional distribution shifts
  • C.3 Comparison with alternative methods
  • C.3.1 Output-space ensembles
  • C.3.2 Comparison to exponential moving averages
  • C.3.3 Additional comparisons when fine-tuning a linear classifier
  • C.4 Changes in data augmentation
  • C.5 Accuracy improvements on reference datasets
  • C.6 Robustness across scales of pre-training compute
  • C.7 WiSE-FT and additional models
  • C.7.1 ALIGN
  • C.7.2 JFT pre-training
  • C.7.3 BASIC
  • C.8 Ensembling zero-shot CLIP with independently trained models
  • D Experimental details
  • D.1 CLIP zero-shot
  • D.2 End-to-end fine-tuning
  • D.3 Fine-tuning a linear classifier
  • D.4 ObjectNet
  • E Diversity measures
  • F When do weight-space ensembles approximate output-space ensembles?

Knowls

  1. Knowl 1 — WiSE-FT interpolates zero-shot and fine-tuned weights

    model/method

    WiSE-FT first fine-tunes a zero-shot model on the target dataset, then forms a single model by element-wise linear interpolation between the zero-shot and fine-tuned parameter vectors. For input image xx, let f(x;θ)f(x;\theta) be the model’s class-score vector, let θ0\theta_0 be the zero-shot parameters, and let θ1\theta_1 be the parameters after standard fine-tuning. WiSE-FT predicts with

    f(x;θα),θα=(1−α)θ0+αθ1,α∈[0,1].f(x;\theta_\alpha),\qquad \theta_\alpha=(1-\alpha)\theta_0+\alpha\theta_1,\qquad \alpha\in[0,1].

    The parameter vectors must correspond to the same model architecture and parameter ordering. The paper recommends α=0.5\alpha=0.5 when there is no domain knowledge for choosing a mixing value. Because inference uses one interpolated parameter set rather than two model passes, WiSE-FT adds no inference computation; evaluating different α\alpha values also requires no additional fine-tuning.

  2. Knowl 2 — WiSE-FT improves CLIP accuracy across ImageNet and five natural shifts

    data/table

    For CLIP ViT-L/14@336px, end-to-end WiSE-FT with α=0.5\alpha=0.5 raises ImageNet accuracy from 86.2% for standard end-to-end fine-tuning to 86.8%, while raising mean accuracy on five ImageNet-derived shifts from 68.6% to 76.9%. With an optimally selected mixing coefficient, these values are 87.1% and 77.4%. The five shift datasets are ImageNet-V2, ImageNet-R, ImageNet Sketch, ObjectNet, and ImageNet-A. The table reports top-1 accuracy in percent; “Avg. shifts” is the mean across those five datasets, and “Avg. ref., shifts” is the mean of ImageNet accuracy and Avg. shifts. “LC” means fine-tuning only a linear classifier; “E2E” means end-to-end fine-tuning. The reported optimal-α\alpha rows use a single coefficient selected to maximize the corresponding average column.

    Method ImageNet IN-V2 IN-R IN-Sketch ObjectNet IN-A Avg. shifts Avg. ref., shifts
    Zero-shot [79] 76.2 70.1 88.9 60.2 70.0 77.2 73.3 74.8
    Fine-tuned LC [79] 85.4 75.9 84.2 57.4 66.2 75.3 71.8 78.6
    Zero-shot (PyTorch) 76.6 70.5 89.0 60.9 69.1 77.7 73.4 75.0
    Fine-tuned LC (ours) 85.2 75.8 85.3 58.7 67.2 76.1 72.6 78.9
    Fine-tuned E2E (ours) 86.2 76.8 79.8 57.9 63.3 65.4 68.6 77.4
    WiSE-FT LC, α=0.5\alpha=0.5 83.7 76.3 89.6 63.0 70.7 79.7 75.9 79.8
    WiSE-FT LC, optimal α\alpha 85.3 76.9 89.8 63.0 70.7 79.7 75.9 80.2
    WiSE-FT E2E, α=0.5\alpha=0.5 86.8 79.5 89.4 64.7 71.1 79.9 76.9 81.8
    WiSE-FT E2E, optimal α\alpha 87.1 79.5 90.3 65.0 72.1 81.0 77.4 81.9
  3. Knowl 3 — WiSE-FT improves performance on six additional distribution shifts

    empirical result

    With α=0.5\alpha=0.5, WiSE-FT improved accuracy relative to the corresponding fine-tuned model on six additional shifts: by 3.5 percentage points (pp) on WILDS-FMoW, 6.2 pp on WILDS-iWildCam, 1.7 pp on CIFAR-10.1, 2.1 pp on CIFAR-10.2, 9.0 pp on ImageNet-Vid-Robust, and 23.2 pp on YTBB-Robust. These evaluations cover geographic shifts in satellite imagery and wildlife recognition, dataset reproductions, and temporal shifts in video. Reference-distribution accuracy decreased by at most 0.3 pp, and often increased. On the WILDS shifts, the zero-shot model initially had less than 30% accuracy; WiSE-FT still improved on fine-tuning.

  4. Knowl 4 — Weight interpolation mitigates fine-tuning hyperparameter brittleness

    empirical result

    For CLIP fine-tuned on ImageNet, modest changes in learning rate or training duration substantially changed accuracy on shifted distributions, even when ImageNet accuracy changed little. For example, after 10 epochs, learning rates of 3×10−53\times10^{-5} and 3×10−63\times10^{-6} produced only a 0.3 pp difference in ImageNet accuracy but up to an 8 pp difference in average accuracy across five shifts. Increasing the learning rate from 10−710^{-7} to 3×10−53\times10^{-5} improved ImageNet accuracy by 5 pp while reducing shifted-distribution accuracy by 8 pp. In the tested configurations, varying the WiSE-FT mixing coefficient could reach a reference-versus-shift accuracy frontier that no tested single hyperparameter configuration matched or exceeded. Thus, selecting among interpolated models by changing α\alpha can address the observed trade-off without training new models.

  5. Knowl 5 — WiSE-FT raises fine-tuned accuracy across seven transfer datasets

    data/table

    The authors report that WiSE-FT improves accuracy over standard end-to-end fine-tuning on all seven datasets shown, including ImageNet and six transfer-learning datasets. At α=0.5\alpha=0.5, absolute gains range from 0.6 to 2.8 pp; with an optimally selected α\alpha, they range from 0.8 to 3.3 pp. Across these datasets, the paper reports relative error reductions of 4%–49% compared with standard fine-tuning. Each entry is top-1 accuracy in percent, with the gain over standard fine-tuning in parentheses.

    Method ImageNet CIFAR-10 CIFAR-100 Cars DTD SUN397 Food-101
    Standard fine-tuning 86.2 98.6 92.2 91.6 81.9 80.7 94.4
    WiSE-FT, α=0.5\alpha=0.5 86.8 (+0.6) 99.3 (+0.7) 93.3 (+1.1) 93.3 (+1.7) 84.6 (+2.8) 83.2 (+2.5) 96.1 (+1.6)
    WiSE-FT, optimal α\alpha 87.1 (+0.9) 99.5 (+0.8) 93.4 (+1.2) 93.6 (+2.0) 85.2 (+3.3) 83.3 (+2.6) 96.2 (+1.8)
  6. Knowl 6 — Classifier-only WiSE-FT is an output ensemble; full-model interpolation stays accurate

    theoretical result

    When fine-tuning changes only a linear classifier and leaves the image encoder fixed, WiSE-FT is exactly equivalent to averaging the two models’ class scores. If g(x)g(x) is the fixed image-feature vector for image xx, and W0W_0 and W1W_1 are the zero-shot and fine-tuned classifier matrices, respectively, then the interpolated classifier produces

    g(x)⊤((1−α)W0+αW1)=(1−α)g(x)⊤W0+αg(x)⊤W1.g(x)^\top\big((1-\alpha)W_0+\alpha W_1\big)=(1-\alpha)g(x)^\top W_0+\alpha g(x)^\top W_1.

    The paper’s experiments identify complementary predictions from the two models as a likely contributor to this benefit. For end-to-end fine-tuning, score averaging is not generally equivalent to parameter interpolation; nevertheless, the authors observe that accuracy remains high along the linear parameter path between zero-shot and fine-tuned models. They relate this observation to linear mode connectivity, while leaving the exact loss-landscape geometry unresolved.

  7. Knowl 7 — WiSE-FT also benefits non-CLIP zero-shot models

    empirical result

    The reported gains extend beyond CLIP. For BASIC-L, WiSE-FT with α=0.5\alpha=0.5 improves average accuracy across five ImageNet-derived shifts by more than 7 pp and improves ImageNet accuracy by 0.4 pp relative to the fine-tuned model; that model was fine-tuned with a contrastive loss using half of ImageNet’s training data. For a ViT-H/14 model pre-trained on JFT-300M, α=0.8\alpha=0.8 improves shifted-distribution accuracy by 2.2 pp over fine-tuning while keeping ImageNet accuracy within 0.2 pp of the fine-tuned model. The paper also reports similar trends for ALIGN, a separately pre-trained vision-language model. The JFT model’s zero-shot classifier used a manually constructed correspondence between its classes and ImageNet classes.

  8. Knowl 8 — Evaluation measures accuracy on a reference distribution and natural shifts

    experimental setup

    The experiments compare accuracy on a reference distribution DrefD_{\mathrm{ref}}, whose training data are used for fine-tuning, with accuracy on a related shifted distribution DshiftD_{\mathrm{shift}}. For a classifier ff, write these test accuracies as Acc⁡ref(f)\operatorname{Acc}_{\mathrm{ref}}(f) and Acc⁡shift(f)\operatorname{Acc}_{\mathrm{shift}}(f). The five principal natural shifts use ImageNet as the reference: ImageNet-V2 (a reproduction of its test set), ImageNet-R (renditions such as paintings and sculptures for 200 classes), ImageNet Sketch, ObjectNet (with 113 overlapping classes), and ImageNet-A (natural images misclassified by ResNet-50 for 200 classes). The paper also uses the effective-robustness measure ρ(f)=Acc⁡shift(f)−β(Acc⁡ref(f))\rho(f)=\operatorname{Acc}_{\mathrm{shift}}(f)-\beta(\operatorname{Acc}_{\mathrm{ref}}(f)), where β\beta is the baseline expected shifted accuracy at a given reference accuracy for models trained on the reference training set. Positive ρ\rho means performance under shift exceeds that baseline.

  9. Knowl 9 — The study is limited to image classification and leaves task-specific mixing open

    limitation

    The experiments investigate image classification, so the paper does not establish that WiSE-FT works in other domains such as object detection or natural language processing. Although α=0.5\alpha=0.5 gives good overall performance in the reported experiments, the best mixing coefficient for a particular target distribution remains an open selection problem.

Coverage note — Secondary comparisons—including EMA, distillation, extra regularization, prompt learning, stronger augmentation, and additional backbone sweeps—are omitted because they are supporting experiments rather than separate load-bearing contributions; the central method, principal results, broader applicability, and stated limitations are represented.

References

  1. 1.Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta. Better fine-tuning by reducing representational collapse. In International Conference on Learning Representations (ICLR), 2021. https://openreview.net/forum?id=OQ08SN70M1V. 7
  2. 2.Michael A Alcorn, Qi Li, Zhitao Gong, Chengfei Wang, Long Mai, Wei-Shinn Ku, and Anh Nguyen. Strike (with) a pose: Neural networks are easily fooled by strange poses of familiar objects. In Conference on Computer Vision and Pattern Recognition (CVPR), 2019. https://arxiv.org/abs/1811.11553. 3, 7
  3. 3.Anders Andreassen, Yasaman Bahri, Behnam Neyshabur, and Rebecca Roelofs. The evolution of out-of-distribution robustness throughout fine-tuning, 2021. https://arxiv.org/abs/2106.15831. 1, 7
  4. 4.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Advances in Neural Information Processing Systems (NeurIPS), 2019. 3, 4, 7
  5. 5.Eric Bauer and Ron Kohavi. An empirical comparison of voting classification algorithms: Bagging, boosting, and variants. Machine learning, 1999. https://link.springer.com/article/10.1023/A:1007515423169. 7
  6. 6.Sara Beery, Arushi Agarwal, Elijah Cole, and Vighnesh Birodkar. The iwildcam 2021 competition dataset. In Conference on Computer Vision and Pattern Recognition (CVPR) FGVC8 Workshop, 2021. https://arxiv.org/abs/2105.03494. 3, 5, 21, 22
  7. 7.Battista Biggio, Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, and Fabio Roli. Evasion attacks against machine learning at test time. In Joint European conference on machine learning and knowledge discovery in databases, 2013. https://arxiv.org/abs/1708.06131. 3
  8. 8.Battista Biggio and Fabio Roli. Wild patterns: Ten years after the rise of adversarial machine learning. Pattern Recognition, 2018. https://arxiv.org/abs/1712.03141. 3
  9. 9.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models, 2021. https://arxiv.org/abs/2108.07258. 1
  10. 10.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In European Conference on Computer Vision (ECCV), 2014. https://data.vision.ee.ethz.ch/cvl/datasets_extra/food-101/. 3, 25, 26, 27
  11. 11.Leo Breiman. Bagging predictors. Machine learning, 1996. https://link.springer.com/article/10.1007/BF00058655. 5, 7
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, et al. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS), 2020. https://arxiv.org/abs/2005.14165. 7, 8
  13. 13.Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee. Functional map of the world. In Conference on Computer Vision and Pattern Recognition (CVPR), 2018. https://arxiv.org/abs/1711.07846. 3, 5, 21, 22
  14. 14.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Conference on Computer Vision and Pattern Recognition (CVPR), 2014. https://arxiv.org/abs/1311.3618. 3, 25, 26, 27
  15. 15.Jeremy Cohen, Elan Rosenfeld, and Zico Kolter. Certified adversarial robustness via randomized smoothing. In International Conference on Machine Learning (ICML), 2019. https://arxiv.org/abs/1902.02918. 17
  16. 16.Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D Hoffman, et al. Underspecification presents challenges for credibility in modern machine learning, 2020. https://arxiv.org/abs/2011.03395. 7
  17. 17.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Conference on Computer Vision and Pattern Recognition, 2009. https://ieeexplore.ieee.org/document/5206848. 3, 4, 27
  18. 18.Karan Desai and Justin Johnson. Virtex: Learning visual representations from textual annotations. In Conference on Computer Vision and Pattern Recognition (CVPR), 2021. https://arxiv.org/abs/2006.06666. 7
  19. 19.Terrance DeVries and Graham W Taylor. Improved regularization of convolutional neural networks with cutout, 2017. https://arxiv.org/abs/1708.04552. 17
  20. 20.Thomas G Dietterich. Ensemble methods in machine learning. In International workshop on multiple classifier systems, 2000. https://link.springer.com/chapter/10.1007/3-540-45014-9_1. 5, 7
  21. 21.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021. https://arxiv.org/abs/2010.11929. 3, 4, 5, 7, 8, 16, 30, 33
  22. 22.Logan Engstrom, Brandon Tran, Dimitris Tsipras, Ludwig Schmidt, and Aleksander Madry. Exploring the landscape of spatial robustness. In International Conference on Machine Learning (ICML), 2019. https://arxiv.org/abs/1712.02779. 17
  23. 23.Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. Robust physical-world attacks on deep learning visual classification. In Conference on Computer Vision and Pattern Recognition (CVPR), 2018. https://arxiv.org/abs/1707.08945. 7
  24. 24.Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M Roy, and Surya Ganguli. Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel. In Advances in Neural Information Processing Systems (NeurIPS), 2020. https://arxiv.org/abs/2010.15110. 44
  25. 25.Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning (ICML), 2020. https://arxiv.org/abs/1912.05671. 3, 5, 8, 15
  26. 26.Yoav Freund and Robert E Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 1997. https://www.sciencedirect.com/science/article/pii/S002200009791504X. 5, 7
  27. 27.Jerome Friedman, Trevor Hastie, Robert Tibshirani, et al. The elements of statistical learning. Springer series in statistics New York, 2001. 7
  28. 28.Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations (ICLR), 2018. https://arxiv.org/abs/1811.12231. 17
  29. 29.Robert Geirhos, Carlos R Medina Temme, Jonas Rauber, Heiko H Schutt, Matthias Bethge, and Felix A Wichmann. Generalisation in humans and deep neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2018. https://arxiv.org/abs/1808.08750. 3, 7
  30. 30.Raphael Gontijo-Lopes, Yann Dauphin, and Ekin D Cubuk. No one representation to rule them all: Overlapping features of training methods, 2021. https://arxiv.org/abs/2007.01434. 14
  31. 31.Raphael Gontijo-Lopes, Yann Dauphin, and Ekin D. Cubuk. No one representation to rule them all: overlapping features of training methods, 2021. https://arxiv.org/abs/2110.12899. 16
  32. 32.Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe. Qualitatively characterizing neural network optimization problems. In International Conference on Learning Representations (ICLR), 2014. https://arxiv.org/abs/1412.6544. 8
  33. 33.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In International Conference on Machine Learning (ICML), 2017. https://arxiv.org/abs/1706.04599. 14
  34. 34.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. https://arxiv.org/abs/1512.03385. 4
  35. 35.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. International Conference on Computer Vision (ICCV), 2021. https://arxiv.org/abs/2006.16241. 3, 4, 7
  36. 36.Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. International Conference on Learning Representations (ICLR), 2019. https://arxiv.org/abs/1903.12261. 3, 7
  37. 37.Dan Hendrycks, Norman Mu, Ekin D. Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. AugMix: A simple data processing method to improve robustness and uncertainty. In International Conference on Learning Representations (ICLR), 2020. https://arxiv.org/abs/1912.02781. 17
  38. 38.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. Conference on Computer Vision and Pattern Recognition (CVPR), 2021. https://arxiv.org/abs/1907.07174. 3, 4, 7
  39. 39.John Hewitt, Xiang Lisa Li, Sang Michael Xie, Benjamin Newman, and Percy Liang. Ensembles and cocktails: Robust finetuning for natural language generation. In NeurIPS 2021 Workshop on Distribution Shifts, 2021. https://openreview.net/forum?id=qXucB21w1C3. 15, 16
  40. 40.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. In Advances in Neural Information Processing Systems (NeurIPS) Deep Learning Workshop, 2015. https://arxiv.org/abs/1503.02531. 24
  41. 41.Tin Kam Ho. The random subspace method for constructing decision forests. IEEE transactions on pattern analysis and machine intelligence, 1998. https://ieeexplore.ieee.org/document/709601. 40
  42. 42.Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. Averaging weights leads to wider optima and better generalization. In Conference on Uncertainty in Artificial Intelligence (UAI), 2018. https://arxiv.org/abs/1803.05407. 4, 5, 8, 15
  43. 43.Arthur Jacot, Franck Gabriel, and Clement Hongler. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in Neural Information Processing Systems (NeurIPS), 2018. https://arxiv.org/abs/1806.07572. 44
  44. 44.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2102.05918. 1, 3, 4, 5, 7, 8, 16, 30, 31
  45. 45.Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. In Association for Computational Linguistics (ACL), 2019. https://arxiv.org/abs/1911.03437. 7
  46. 46.James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences (PNAS), 2017. https://arxiv.org/abs/1612.00796. 7
  47. 47.Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A. Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang. WILDS: A benchmark of in-the-wild distribution shifts. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2012.07421. 3, 5, 7, 21, 22
  48. 48.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. In European Conference on Computer Vision (ECCV), 2020. https://arxiv.org/abs/1912.11370. 7
  49. 49.Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. Similarity of neural network representations revisited. In International Conference on Machine Learning (ICML), 2019. https://arxiv.org/abs/1905.00414. 14, 41
  50. 50.Simon Kornblith, Jonathon Shlens, and Quoc V Le. Do better imagenet models transfer better? In Conference on Computer Vision and Pattern Recognition (CVPR), 2019. https://arxiv.org/abs/1805.08974. 25, 26
  51. 51.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In International Conference on Computer Vision (ICCV) Workshops, 2013. https://ieeexplore.ieee.org/document/6755945. 3, 25, 26, 27
  52. 52.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images, 2009. https://www.cs.toronto.edu/~kriz/learning-features-2009-TR.pdf. 3, 5, 21, 25, 26, 27
  53. 53.Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. Fine-tuning distorts pretrained features and underperforms out-of-distribution, 2021. https://openreview.net/forum?id=UYneFzXSJWh. 15
  54. 54.Ananya Kumar, Aditi Raghunathan, Tengyu Ma, and Percy Liang. Calibrated ensembles: A simple way to mitigate ID-OOD accuracy tradeoffs. In NeurIPS 2021 Workshop on Distribution Shifts, 2021. https://openreview.net/forum?id=dmDE-9e9F_x. 15
  55. 55.Ludmila I Kuncheva and Christopher J Whitaker. Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy. Machine learning, 2003. https://doi.org/10.1023/A:1022859003006. 14
  56. 56.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In Advances in Neural Information Processing Systems (NeurIPS), 2017. https://arxiv.org/abs/1612.01474. 7, 8
  57. 57.Hao Li, Pratik Chaudhari, Hao Yang, Michael Lam, Avinash Ravichandran, Rahul Bhotika, and Stefano Soatto. Rethinking the hyperparameters for fine-tuning. In International Conference on Learning Representations (ICLR), 2020. https://arxiv.org/abs/2002.11770. 7
  58. 58.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations (ICLR), 2016. https://arxiv.org/abs/1608.03983. 23, 39
  59. 59.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019. https://openreview.net/forum?id=Bkg6RiCqY7. 6, 23, 39
  60. 60.Shangyun Lu, Bradley Nott, Aaron Olson, Alberto Todeschini, Hossein Vahabi, Yair Carmon, and Ludwig Schmidt. Harder or different? a closer look at distribution shift in dataset reproduction. In International Conference on Machine Learning (ICML) Workshop on Uncertainty and Robustness in Deep Learning, 2020. http://www.gatsby.ucl.ac.uk/~balaji/udl2020/accepted-papers/UDL2020-paper-101.pdf. 3, 5, 21, 22
  61. 61.Ekdeep Singh Lubana, Puja Trivedi, Danai Koutra, and Robert P. Dick. How do quadratic regularizers prevent catastrophic forgetting: The role of interpolation, 2021. https://arxiv.org/abs/2102.02805. 7
  62. 62.James Lucas, Juhan Bae, Michael R Zhang, Stanislav Fort, Richard Zemel, and Roger Grosse. Analyzing monotonic linear interpolation in neural network loss landscapes. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2104.11044. 8
  63. 63.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations (ICLR), 2017. https://arxiv.org/abs/1706.06083. 7, 17
  64. 64.Michael Matena and Colin Raffel. Merging models with fisher-weighted averaging, 2021. https://arxiv.org/abs/2111.09832. 16
  65. 65.Michael McCloskey and Neal J. Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. Psychology of Learning and Motivation, 1989. https://www.sciencedirect.com/science/article/pii/S0079742108605368. 7
  66. 66.Mary L McHugh. Interrater reliability: the kappa statistic. Biochemia medica, 2012. 40
  67. 67.John Miller, Karl Krauth, Benjamin Recht, and Ludwig Schmidt. The effect of natural distribution shift on question answering models. In International Conference on Machine Learning (ICML), 2020. https://arxiv.org/abs/2004.14444. 7
  68. 68.John P Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa, Pang Wei Koh, Vaishaal Shankar, Percy Liang, Yair Carmon, and Ludwig Schmidt. Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2107.04649. 1, 4, 7, 17
  69. 69.Rafael Muller, Simon Kornblith, and Geoffrey Hinton. When does label smoothing help? In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1906.02629. 23, 39
  70. 70.Basil Mustafa, Carlos Riquelme, Joan Puigcerver, Andre Susano Pinto, Daniel Keysers, and Neil Houlsby. Deep ensembles for low-data transfer learning, 2020. https://arxiv.org/abs/2010.06866. 8
  71. 71.Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang. What is being transferred in transfer learning? In Advances in Neural Information Processing Systems (NeurIPS), 2020. https://arxiv.org/abs/2008.11687. 3, 4, 5, 15
  72. 72.Alex Nichol, Joshua Achiam, and John Schulman. On first-order meta-learning algorithms, 2018. https://arxiv.org/abs/1803.02999. 8
  73. 73.Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua V Dillon, Balaji Lakshminarayanan, and Jasper Snoek. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1906.02530. 8
  74. 74.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1912.01703. 23, 39
  75. 75.Hieu Pham, Zihang Dai, Golnaz Ghiasi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, and Quoc V. Le. Combined scaling for zero-shot transfer learning, 2021. https://arxiv.org/abs/2111.10050. 1, 3, 4, 5, 7, 8, 16, 30, 34, 35
  76. 76.Boris Teodorovich Polyak. New method of stochastic approximation type. Automation and remote control, 1990. 2
  77. 77.Boris T Polyak and Anatoli B Juditsky. Acceleration of stochastic approximation by averaging. SIAM journal on control and optimization, 1992. https://epubs.siam.org/doi/abs/10.1137/0330046?journalCode=sjcodc. 8
  78. 78.Joaquin Quinonero-Candela, Masashi Sugiyama, Neil D Lawrence, and Anton Schwaighofer. Dataset shift in machine learning. Mit Press, 2009. 7
  79. 79.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2103.00020. 1, 3, 4, 5, 6, 7, 8, 24, 25, 30, 39
  80. 80.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language Models are Unsupervised Multitask Learners, 2019. https://openai.com/blog/better-language-models/. 3, 7
  81. 81.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do ImageNet classifiers generalize to ImageNet? In International Conference on Machine Learning (ICML), 2019. https://arxiv.org/abs/1902.10811. 3, 4, 5, 21, 22
  82. 82.David Ruppert. Efficient estimations from a slowly convergent robbins-monro process, 1988. https://ecommons.cornell.edu/handle/1813/8664. 2, 8
  83. 83.Hadi Salman, Greg Yang, Jerry Li, Pengchuan Zhang, Huan Zhang, Ilya Razenshteyn, and Sebastien Bubeck. Provably robust deep learning via adversarially trained smoothed classifiers. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1906.04584. 17
  84. 84.Mert Bulent Sariyildiz, Julien Perez, and Diane Larlus. Learning visual representations with caption annotations. In European Conference on Computer Vision (ECCV), 2020. https://arxiv.org/abs/2008.01392. 7
  85. 85.Ali Shafahi, Mahyar Najibi, Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein. Adversarial training for free! In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1904.12843. 17
  86. 86.Vaishaal Shankar, Achal Dave, Rebecca Roelofs, Deva Ramanan, Benjamin Recht, and Ludwig Schmidt. Do image classifiers generalize across time?, 2019. https://arxiv.org/abs/1906.02168. 3, 21, 22
  87. 87.Vaishaal Shankar, Rebecca Roelofs, Horia Mania, Alex Fang, Benjamin Recht, and Ludwig Schmidt. Evaluating machine accuracy on imagenet. In International Conference on Machine Learning (ICML), 2020. http://proceedings.mlr.press/v119/shankar20c/shankar20c.pdf. 3, 5
  88. 88.Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson. Cnn features off-the-shelf: an astounding baseline for recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2014. https://arxiv.org/abs/1403.6382. 7
  89. 89.David B Skalak et al. The sources of increased accuracy for two proposed boosting algorithms. In American Association for Artificial Intelligence (AAAI), Integrating Multiple Learned Models Workshop, 1996. https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.40.2269&rep=rep1&type=pdf. 40
  90. 90.Asa Cooper Stickland and Iain Murray. Diverse ensembles improve calibration. In International Conference on Machine Learning (ICML) Workshop on Uncertainty and Robustness in Deep Learning, 2020. https://arxiv.org/abs/2007.04206. 8
  91. 91.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In International Conference on Computer Vision (ICCV), 2017. https://arxiv.org/abs/1707.02968. 5, 7, 8, 33
  92. 92.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. https://arxiv.org/abs/1912.04838. 7
  93. 93.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Conference on Computer Vision and Pattern Recognition (CVPR), 2016. https://arxiv.org/abs/1512.00567. 7, 8, 23
  94. 94.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning (ICML), 2019. https://proceedings.mlr.press/v97/tan19a/tan19a.pdf. 38
  95. 95.Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt. Measuring robustness to natural distribution shifts in image classification. In Advances in Neural Information Processing Systems (NeurIPS), 2020. https://arxiv.org/abs/2007.00644. 1, 2, 3, 4, 7, 17, 18
  96. 96.Antonio Torralba and Alexei A Efros. Unbiased look at dataset bias. In Conference on Computer Vision and Pattern Recognition (CVPR), 2011. https://people.csail.mit.edu/torralba/publications/datasets_cvpr11.pdf. 7
  97. 97.Florian Tramer, Alexey Kurakin, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel. Ensemble adversarial training: Attacks and defenses. In International Conference on Learning Representations (ICLR), 2017. https://arxiv.org/abs/1705.07204. 7
  98. 98.Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1905.13549. 3, 4, 7
  99. 99.Ross Wightman. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019. 23, 25, 26, 39
  100. 100.Mitchell Wortsman, Maxwell C Horton, Carlos Guestrin, Ali Farhadi, and Mohammad Rastegari. Learning neural network subspaces. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2102.10472. 15
  101. 101.Jianxiong Xiao, Krista A Ehinger, James Hays, Antonio Torralba, and Aude Oliva. Sun database: Exploring a large collection of scene categories. International Journal of Computer Vision, 2016. https://link.springer.com/article/10.1007/s11263-014-0748-y. 3, 25, 26, 27
  102. 102.Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le. Self-training with noisy student improves imagenet classification. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. https://arxiv.org/abs/1911.04252. 38
  103. 103.LI Xuhong, Yves Grandvalet, and Franck Davoine. Explicit inductive bias for transfer learning with convolutional networks. In International Conference on Machine Learning (ICML), 2018. https://arxiv.org/abs/1802.01483. 7
  104. 104.Chhavi Yadav and Leon Bottou. Cold case: The lost mnist digits. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1905.10498. 7
  105. 105.I Zeki Yalniz, Herve Jegou, Kan Chen, Manohar Paluri, and Dhruv Mahajan. Billion-scale semi-supervised learning for image classification, 2019. https://arxiv.org/abs/1905.00546. 7
  106. 106.Friedemann Zenke, Ben Poole, and Surya Ganguli. Continual learning through synaptic intelligence. In International Conference on Machine Learning (ICML), 2017. https://arxiv.org/abs/1703.04200. 7
  107. 107.Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. Lit: Zero-shot transfer with locked-image text tuning, 2021. https://arxiv.org/abs/2111.07991. 7
  108. 108.Michael R Zhang, James Lucas, Geoffrey Hinton, and Jimmy Ba. Lookahead optimizer: k steps forward, 1 step back. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1907.08610. 8
  109. 109.Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz. Contrastive learning of medical visual representations from paired images and text, 2020. https://arxiv.org/abs/2010.00747. 7
  110. 110.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models, 2021. https://arxiv.org/abs/2109.01134. 7, 16, 24, 25
  111. 111.Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. Freelb: Enhanced adversarial training for natural language understanding. In International Conference on Learning Representations (ICLR), 2020. https://arxiv.org/abs/1909.11764. 7

Citation

MLA
Wortsman, M., et al. “Robust Fine-tuning of Zero-shot Models”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 7949–61, https://doi.org/10.1109/CVPR52688.2022.00780.
APA
Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Lopes, R. G., Hajishirzi, H., Farhadi, A., Namkoong, H., & Schmidt, L. (2022). Robust fine-tuning of zero-shot models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7949–7961. https://doi.org/10.1109/CVPR52688.2022.00780
Chicago
Wortsman, M., G. Ilharco, J. W. Kim, et al. 2022. “Robust Fine-tuning of Zero-shot Models”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7949–61. https://doi.org/10.1109/CVPR52688.2022.00780.
Harvard
Wortsman, M. et al. (2022) “Robust fine-tuning of zero-shot models”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 7949–7961. Available at: https://doi.org/10.1109/CVPR52688.2022.00780.
Vancouver
1. Wortsman M, Ilharco G, Kim JW, et al (2022) Robust fine-tuning of zero-shot models. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 7949–7961

BibTeX

@inproceedings{Wortsman_2022, title={Robust fine-tuning of zero-shot models}, url={http://dx.doi.org/10.1109/CVPR52688.2022.00780}, DOI={10.1109/cvpr52688.2022.00780}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Wortsman, Mitchell and Ilharco, Gabriel and Kim, Jong Wook and Li, Mike and Kornblith, Simon and Roelofs, Rebecca and Lopes, Raphael Gontijo and Hajishirzi, Hannaneh and Farhadi, Ali and Namkoong, Hongseok and Schmidt, Ludwig}, year={2022}, month=June, pages={7949–7961} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE