Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Mitchell WortsmanGabriel IlharcoSamir Yitzhak GadreRebecca RoelofsRaphael Gontijo-LopesAri S. MorcosHongseok NamkoongAli FarhadiYair CarmonSimon Kornblith

article2022ICML1,739 citations

Introduces model soups, a method of averaging the weights of multiple fine-tuned models that yields ensemble-level accuracy and out-of-distribution generalization without adding any computational or memory overhead at inference time.

Listen

In modern machine learning workflows, organizations frequently fine-tune large pre-trained foundation models across a broad range of training configurations to maximize task performance. Standard practice retains only the single highest-performing model on a validation dataset and discards the rest. While combining predictions from multiple models into an ensemble reliably improves accuracy and robustness, it multiplies inference runtime and computational memory costs proportionally to the number of models used, making real-time deployment expensive or impractical.

The article demonstrates that averaging the parameter weights of multiple models fine-tuned from the same shared pre-trained initialization—a method called "model soups"—improves predictive accuracy and out-of-distribution robustness without incurring any additional inference latency, compute, or memory overhead relative to a single model.

To evaluate this approach, the researchers conducted extensive empirical evaluations across vision architectures (including CLIP, ALIGN, BASIC, and ViT-G) and language transformer models (BERT and T5) across varied hyperparameter sweeps, data augmentations, and optimizers. The evaluation benchmarked standard validation performance as well as generalization under real-world distribution shifts. The core technique introduced, the greedy soup, sorts fine-tuned models by validation accuracy and sequentially blends their weights into an aggregate model only if the combination improves validation accuracy.

The primary finding is that greedy model soups consistently outperform the single best model identified during hyperparameter sweeps. Notably, applying a greedy soup to a large vision transformer (ViT-G) achieved a state-of-the-art 90.94% top-1 accuracy on ImageNet while utilizing 25% fewer floating-point operations during inference than the previous leading architecture. Second, model soups significantly enhance robustness against distribution shifts, often matching or exceeding the performance of traditional output ensembles without requiring multiple model evaluation passes. Third, the benefits generalize across modalities, delivering performance gains on multiple text classification tasks and cross-dataset zero-shot transfers.

These findings have immediate cost and operational implications for machine learning deployment. Practitioners can achieve ensemble-level accuracy and robustness directly from the intermediate outputs of standard hyperparameter searches at zero additional inference and serving cost. Unlike conventional ensembling, model soups require no architectural modifications or latency trade-offs in production.

Organizations fine-tuning foundation models should integrate greedy weight averaging into their deployment pipelines as a standard post-processing step rather than selecting a single checkpoint. For future development, teams can design hyperparameter sweeps to maximize model diversity, such as varying learning rates and augmentation strategies, to maximize the benefits of weight interpolation.

Decision-makers should note key boundaries: model soups require fine-tuned models to share the exact same pre-trained initialization, and performance gains are less substantial when pre-training data is small or homogeneous. Furthermore, while model soups reliably boost classification accuracy, they do not replicate the predictive calibration and uncertainty estimation improvements provided by traditional ensembles.

Cover for Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Abstract

The conventional recipe for maximizing model accuracy is to (1) train multiple models with various hyperparameters and (2) pick the individual model which performs best on a held-out validation set, discarding the remainder. In this paper, we revisit the second step of this procedure in the context of fine-tuning large pre-trained models, where fine-tuned models often appear to lie in a single low error basin. We show that averaging the weights of multiple models fine-tuned with different hyperparameter configurations often improves accuracy and robustness. Unlike a conventional ensemble, we may average many models without incurring any additional inference or memory costs -- we call the results "model soups." When fine-tuning large pre-trained models such as CLIP, ALIGN, and a ViT-G pre-trained on JFT, our soup recipe provides significant improvements over the best model in a hyperparameter sweep on ImageNet. The resulting ViT-G model, which attains 90.94% top-1 accuracy on ImageNet, achieved a new state of the art. Furthermore, we show that the model soup approach extends to multiple image classification and natural language processing tasks, improves out-of-distribution performance, and improves zero-shot performance on new downstream tasks. Finally, we analytically relate the performance similarity of weight-averaging and logit-ensembling to flatness of the loss and confidence of the predictions, and validate this relation empirically. Code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 3 Experiments
  • 3.1 Experimental setup
  • 3.2 Intuition and motivation
  • 3.3 Model soups
  • 3.3.1 Fine-tuning CLIP and ALIGN
  • 3.3.2 Fine-tuning a ViT-G model pre-trained on JFT-3B
  • 3.3.3 Fine-tuning on text classification tasks
  • 4 Analytically comparing soups to ensembles
  • 5 Scope and limitations
  • 6 Related work
  • 7 Conclusion
  • References
  • A Overview
  • B Additional figures
  • C BASIC
  • D Robust fine-tuning
  • E Cross-dataset soups
  • F Analysis of 1D hyperparameter grids
  • G Additional fine-tuning and pre-training datasets
  • H Additional grid searches and initializations
  • I Learned soup
  • J Experimental details
  • J.1 Error landscape visualizations
  • J.2 Model soups
  • J.2.1 CLIP experiments
  • J.2.2 ALIGN experiments
  • J.2.3 ViT-G/14 experiments
  • J.3 Cross-dataset soups details
  • J.4 Text classification datasets
  • J.5 Fine-tuning details for text classification tasks
  • K Analytical comparison details
  • K.1 Notation and preliminaries
  • K.2 An exact expression for logit difference
  • K.3 Derivation of approximation
  • K.4 Detailed empirical evaluations
  • L Additional baselines

Knowls

  1. Knowl 1 — Greedy Soup Algorithm for Fine-Tuned Model Weight Averaging

    algorithm

    The Greedy Soup procedure constructs an averaged model by sequentially evaluating candidates from a hyperparameter sweep on a held-out validation set and retaining only those that do not degrade validation performance.

    Input: Potential soup ingredients {θ1,…,θk\theta_1, \dots, \theta_k} sorted in descending order of held-out validation accuracy ValAcc(θi)\text{ValAcc}(\theta_i)
    Output: Averaged parameter vector θS\theta_S
    ingredients ←\leftarrow {θ1\theta_1}
    for i=2i = 2 to kk do
        candidate ←\leftarrow ingredients ∪\cup {θi\theta_i}
        if ValAcc(average(candidate))≥ValAcc(average(ingredients))\text{ValAcc}(\text{average}(\text{candidate})) \ge \text{ValAcc}(\text{average}(\text{ingredients})) then
            ingredients ←\leftarrow candidate
    return $\text{average}(\text{ingredients})

    The algorithm requires k−1k-1 evaluations on the held-out validation set, running in O(k)O(k) time with respect to validation evaluations. Because candidate models are evaluated after sorting by individual performance, the resulting greedy soup is guaranteed to perform at least as well on the validation set as the single best individual fine-tuned model. At inference time, the resulting model requires O(1)O(1) compute and memory relative to a single model.

  2. Knowl 2 — Model Soups via Weight Averaging of Independently Fine-Tuned Models

    definition

    Let θ0∈Rd\theta_0 \in \mathbb{R}^d be a shared pre-trained parameter initialization, and let θi=FineTune(θ0,hi)∈Rd\theta_i = \text{FineTune}(\theta_0, h_i) \in \mathbb{R}^d denote the parameters obtained by fine-tuning θ0\theta_0 with hyperparameter configuration hih_i for i∈{1,…,k}i \in \{1, \dots, k\}. A model soup is a neural network whose parameters θS\theta_S are defined as the arithmetic mean of a selected subset of fine-tuned weights S⊆{1,…,k}S \subseteq \{1, \dots, k\}:

    θS=1∣S∣∑i∈Sθi\theta_S = \frac{1}{|S|} \sum_{i \in S} \theta_i

    A uniform soup averages all fine-tuned configurations (S={1,…,k}S = \{1, \dots, k\}). While a conventional ensemble computes the average of model predictions fens(x)=1k∑i=1kf(x;θi)f_{\text{ens}}(x) = \frac{1}{k}\sum_{i=1}^k f(x; \theta_i) and incurs an O(k)O(k) multiplier on inference compute and memory, a model soup f(x;θS)f(x; \theta_S) requires O(1)O(1) inference compute and memory identical to a standard single network.

  3. Knowl 3 — Analytical Relation Between Model Soup and Logit Ensemble Loss

    theoretical result

    Let θ0,θ1∈Rd\theta_0, \theta_1 \in \mathbb{R}^d be two fine-tuned parameter vectors and let θα=(1−α)θ0+αθ1\theta_\alpha = (1-\alpha)\theta_0 + \alpha\theta_1 with α∈[0,1]\alpha \in [0, 1] denote the weight-averaged model. Let f(x;θ)∈RCf(x; \theta) \in \mathbb{R}^C denote the logit output of a network for CC-class classification, and let fαens(x)=(1−α)f(x;θ0)+αf(x;θ1)f^{\text{ens}}_\alpha(x) = (1-\alpha)f(x; \theta_0) + \alpha f(x; \theta_1) denote the logit ensemble. For cross-entropy loss ℓ(f,y)=log⁡∑y′exp⁡(fy′−fy)\ell(f, y) = \log \sum_{y'} \exp(f_{y'} - f_y) and temperature parameter β>0\beta > 0, define the β\beta-calibrated expected losses as Lαsoup=Ex,y[ℓ(βf(x;θα),y)]\mathcal{L}_\alpha^{\text{soup}} = \mathbb{E}_{x, y}[\ell(\beta f(x; \theta_\alpha), y)] and Lαens=Ex,y[ℓ(βfαens(x),y)]\mathcal{L}_\alpha^{\text{ens}} = \mathbb{E}_{x, y}[\ell(\beta f^{\text{ens}}_\alpha(x), y)].

    Assuming the network logits are approximately quadratic along the linear segment connecting θ0\theta_0 and θ1\theta_1, the expected loss difference between the weight-averaged soup and the logit ensemble is approximated by:

    Lαsoup−Lαens≈α(1−α)2[−d2dα2Lαsoup+β2Ex[VarY∼psftmx(βf(x;θα))[fY(x;θ1)−fY(x;θ0)]]]\mathcal{L}_\alpha^{\text{soup}} - \mathcal{L}_\alpha^{\text{ens}} \approx \frac{\alpha(1-\alpha)}{2} \left[ -\frac{d^2}{d\alpha^2}\mathcal{L}_\alpha^{\text{soup}} + \beta^2 \mathbb{E}_{x} \left[ \text{Var}_{Y \sim p_{\text{sftmx}}(\beta f(x; \theta_\alpha))} \left[ f_Y(x; \theta_1) - f_Y(x; \theta_0) \right] \right] \right]

    where [psftmx(f)]i=efi/∑jefj[p_{\text{sftmx}}(f)]_i = e^{f_i} / \sum_j e^{f_j}.

    The first term implies that negative curvature (flatness/convexity along the interpolation path) favors the weight-averaged soup. The second term always favors the logit ensemble, but vanishes when either the endpoint predictions Δf(x)=f(x;θ1)−f(x;θ0)\Delta f(x) = f(x; \theta_1) - f(x; \theta_0) are small or the soup model produces highly confident predictions (causing the softmax distribution to concentrate on a single class and minimizing the variance).

  4. Knowl 4 — State-of-the-Art ImageNet Classification with ViT-G/14 Greedy Soup

    data/table

    Applying the greedy soup recipe to 58 fine-tuned checkpoints of a ViT-G/14 model pre-trained on JFT-3B selects 14 model checkpoints and establishes a new state of the art on ImageNet top-1 accuracy (90.94%) while outperforming individual models across multiple natural distribution shifts.

    Method Top-1 ReaL Multilabel IN-V2 IN-R IN-Sketch ObjectNet IN-A Avg shifts
    ViT/G-14 (original) 90.45 90.81 – 83.33 – – 70.53 – –
    CoAtNet-7 90.88 – – – – – – – –
    Best model on val set 90.72 91.04 96.94 83.76 95.04 73.16 78.20 91.75 84.38
    Best model on each test set 90.78 91.78 97.29 84.31 95.04 73.73 79.03 92.16 84.68
    Greedy ensemble 90.93 91.29 97.23 84.14 94.85 73.07 77.87 91.69 84.33
    Greedy soup 90.94 91.20 97.17 84.22 95.46 74.23 78.52 92.67 85.02

    The table compares accuracy (%) on the standard ImageNet validation set (Top-1), ReaL, and Multilabel sets, alongside five out-of-distribution test sets: ImageNet-V2 (IN-V2), ImageNet-R (IN-R), ImageNet-Sketch, ObjectNet, and ImageNet-A (IN-A). Greedy soup exceeds the best validation-selected model on 7 out of 8 evaluation benchmarks without requiring the O(k)O(k) computational inference overhead of greedy logit ensembling.

  5. Knowl 5 — Learned Soup Optimization via Validation Loss Minimization

    model/method

    The learned soup formulation replaces the discrete, greedy inclusion step of model souping with continuous optimization of mixing coefficients α∈Rk\alpha \in \mathbb{R}^k and a temperature scaling parameter β∈R\beta \in \mathbb{R} over a held-out validation set {(xj,yj)}j=1n\{(x_j, y_j)\}_{j=1}^n:

    arg⁡min⁡α∈Rk,β∈R∑j=1nℓ(β⋅f(xj,∑i=1kαiθi),yj)\arg\min_{\alpha \in \mathbb{R}^k, \beta \in \mathbb{R}} \sum_{j=1}^n \ell \left( \beta \cdot f\left( x_j, \sum_{i=1}^k \alpha_i \theta_i \right), y_j \right)

    where ℓ\ell is the cross-entropy loss and θi\theta_i are candidate fine-tuned parameter vectors. To enforce that the coefficients lie on the probability simplex (αi≥0\alpha_i \ge 0 and ∑i=1kαi=1\sum_{i=1}^k \alpha_i = 1), α\alpha is parameterized as the output of a softmax function. Optimization is performed using mini-batch gradient descent (e.g., AdamW with a constant learning rate of 0.1 for 3 epochs). In a layer-wise variant, distinct mixing vectors α(l)\alpha^{(l)} are optimized independently for each network layer ll.

  6. Knowl 6 — Performance and Robustness Gains on Vision-Language Models (CLIP and ALIGN)

    empirical result

    When fine-tuning CLIP ViT-B/32 and ALIGN EfficientNet-L2 on ImageNet across diverse hyperparameter configurations (varying learning rates, weight decays, training durations, augmentations, and mixup rates):

    1. For CLIP ViT-B/32, fine-tuning over a 72-model random hyperparameter sweep yields a best individual model achieving 80.38% top-1 ImageNet accuracy and 47.83% average accuracy on five distribution shift datasets (ImageNet-V2, ImageNet-R, ImageNet-Sketch, ObjectNet, ImageNet-A). The greedy soup selects 5 models and attains 81.03% top-1 on ImageNet (+0.65 pp) and 50.75% on distribution shifts (+2.92 pp).
    2. For ALIGN EfficientNet-L2, fine-tuning over a 12-model grid search yields a greedy soup (selecting 5 models) that improves ImageNet top-1 accuracy by 0.5 percentage points over the best individual fine-tuned model and outperforms individual runs under distribution shifts.
    3. The performance of the greedy soup exceeds that of individual models across nearly all intermediate soup sizes.
  7. Knowl 7 — Correlation Between Parameter Angle and Model Soup Accuracy Gains

    empirical result

    For pairs of models θ1,θ2∈Rd\theta_1, \theta_2 \in \mathbb{R}^d fine-tuned from a shared pre-trained initialization θ0\theta_0, the interpolation advantage—defined as the accuracy of the midpoint model minus the average accuracy of the individual endpoints:

    Acc(12θ1+12θ2)−12(Acc(θ1)+Acc(θ2))\text{Acc}\left(\frac{1}{2}\theta_1 + \frac{1}{2}\theta_2\right) - \frac{1}{2}\left(\text{Acc}(\theta_1) + \text{Acc}(\theta_2)\right)

    is positively correlated with the angle ϕ\phi between the parameter displacement vectors θ1−θ0\theta_1 - \theta_0 and θ2−θ0\theta_2 - \theta_0.

    Varying hyperparameter configurations (such as learning rate, data augmentation strength, and random seeds) increases the angle ϕ\phi toward 90∘90^\circ, producing more orthogonal solutions that yield higher interpolation gains upon weight averaging, provided the learning rates remain within the basin of the shared initialization.

  8. Knowl 8 — Cross-Dataset Backbone Soups for Enhanced Zero-Shot Generalization

    empirical result

    Averaging the backbone weights of models fine-tuned on multiple distinct image classification datasets improves zero-shot transfer performance on an unseen held-out task.

    A CLIP ViT-B/32 model is fine-tuned independently on six separate datasets (CIFAR-10, Describable Textures, Food-101, SUN397, Stanford Cars, and ImageNet) while keeping the classification head produced by the CLIP text tower fixed so that all task adaptations are restricted to backbone weights. When the resulting six backbones and the original zero-shot backbone are averaged into a single backbone soup and paired with a zero-shot classification head for CIFAR-100 (a dataset not seen during fine-tuning), the zero-shot top-1 accuracy on CIFAR-100 improves by 6.4 percentage points over the base zero-shot CLIP model.

  9. Knowl 9 — Performance of Greedy Model Soups on NLP Text Classification (GLUE)

    data/table

    Model soups improve validation metrics on natural language processing tasks when fine-tuning transformer architectures over random hyperparameter sweeps on the GLUE benchmark.

    Model Method MRPC (Acc/F1) RTE (Acc) CoLA (MCC) SST-2 (Acc)
    BERT-base Best individual 88.3 61.0 59.1 92.5
    BERT-base Greedy soup 88.3 61.7 59.1 93.0
    T5-base Best individual 91.8 78.3 58.8 94.6
    T5-base Greedy soup 92.4 79.1 60.2 94.7

    Across 32 runs per dataset varying learning rate, batch size, epochs, and random seed, the greedy soup achieves improvements or matches the best individual model on MRPC, RTE, CoLA, and SST-2 for both BERT-base and T5-base. Across all 20 tested combinations of five model scales (BERT-base, BERT-large, T5-small, T5-base, T5-large) and four datasets, greedy soup strictly outperforms the best individual model in 10 configurations and matches it in the remainder.

  10. Knowl 10 — Additivity of Model Souping with Intra-Run Trajectory Averaging

    empirical result

    Model souping across independently fine-tuned runs is additive with intra-run weight averaging techniques such as Exponential Moving Average (EMA) and Stochastic Weight Averaging (SWA).

    When individual fine-tuning runs maintain an EMA of weights (with decay factor β=0.999\beta = 0.999) or apply SWA along their training trajectories, applying the greedy soup procedure across these smoothed checkpoints yields higher top-1 accuracy and out-of-distribution robustness than either intra-run trajectory averaging or inter-run model souping alone. In ViT-G/14 experiments, combining low EMA decay checkpointing with greedy souping provided the highest overall performance.

  11. Knowl 11 — Calibration Deficit of Model Soups Compared to Ensembles

    limitation

    While deep ensembles consistently reduce Expected Calibration Error (ECE) under both in-distribution and out-of-distribution evaluation, weight-averaged model soups do not substantially improve calibration over individual models.

    Evaluating soups and ensembles formed from 20 identically configured CLIP ViT-B/32 fine-tuning runs (varying only random seeds) shows that logit ensembling markedly reduces ECE on ImageNet and distribution shift datasets, whereas uniform model soups maintain ECE values comparable to single fine-tuned models. Pre-calibrating individual models via temperature scaling prior to weight averaging does not resolve this discrepancy.

Coverage note — None was omitted; all primary methods (uniform, greedy, learned soups), analytical loss bounds, main empirical findings across vision (ViT-G, CLIP, ALIGN, BASIC) and NLP (BERT, T5), cross-dataset transfer, trajectory averaging interactions, and calibration limitations are fully covered.

References

  1. 1.Pulkit Agrawal, Ross Girshick, and Jitendra Malik. Analyzing the performance of multilayer neural networks for object recognition. In European conference on computer vision, pages 329–344. Springer, 2014.
  2. 2.Anders Andreassen, Yasaman Bahri, Behnam Neyshabur, and Rebecca Roelofs. The evolution of out-of-distribution robustness throughout fine-tuning, 2021. https://arxiv.org/abs/2106.15831.
  3. 3.Hossein Azizpour, Ali Sharif Razavian, Josephine Sullivan, Atsuto Maki, and Stefan Carlsson. From generic to specific deep representations for visual recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 36–45, 2015.
  4. 4.Hessam Bagherinezhad, Maxwell Horton, Mohammad Rastegari, and Ali Farhadi. Label refinery: Improving imagenet classification through label progression. arXiv preprint arXiv:1805.02641, 2018.
  5. 5.Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. The second pascal recognising textual entailment challenge. In Proc. of the II PASCAL challenge, 2006.
  6. 6.Andrei Barbu, David Mayo, Julian Alverio, William Luo, Christopher Wang, Dan Gutfreund, Josh Tenenbaum, and Boris Katz. Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In Advances in Neural Information Processing Systems (NeurIPS), 2019. URL https://proceedings.neurips.cc/paper/2019/file/97af07a14cacba681feacf3012730892-Paper.pdf.
  7. 7.Eric Bauer and Ron Kohavi. An empirical comparison of voting classification algorithms: Bagging, boosting, and variants. Machine learning, 1999. https://link.springer.com/article/10.1023/A:1007515423169.
  8. 8.Sara Beery, Arushi Agarwal, Elijah Cole, and Vighnesh Birodkar. The iwildcam 2021 competition dataset. In Conference on Computer Vision and Pattern Recognition (CVPR) FGVC8 Workshop, 2021. https://arxiv.org/abs/2105.03494.
  9. 9.Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. The fifth pascal recognizing textual entailment challenge. In TAC, 2009.
  10. 10.Lucas Beyer, Olivier J Hénaff, Alexander Kolesnikov, Xiaohua Zhai, and Aaron van den Oord. Are we done with imagenet? arXiv preprint arXiv:2006.07159, 2020.
  11. 11.Lucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov. Knowledge distillation: A good teacher is patient and consistent, 2021. URL https://arxiv.org/abs/2106.05237.
  12. 12.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models, 2021. https://arxiv.org/abs/2108.07258.
  13. 13.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In European Conference on Computer Vision (ECCV), 2014. https://data.vision.ee.ethz.ch/cvl/datasets_extra/food-101/.
  14. 14.Leo Breiman. Bagging predictors. Machine learning, 1996. https://link.springer.com/article/10.1007/BF00058655.
  15. 15.Rich Caruana, Alexandru Niculescu-Mizil, Geoff Crew, and Alex Ksikes. Ensemble selection from libraries of models. In Proceedings of the twenty-first international conference on Machine learning, page 18, 2004.
  16. 16.Rich Caruana, Art Munson, and Alexandru Niculescu-Mizil. Getting the most out of ensemble selection. In Sixth International Conference on Data Mining (ICDM'06), pages 828–833. IEEE, 2006.
  17. 17.Ken Chatfield, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Return of the devil in the details: Delving deep into convolutional nets. In British Machine Vision Conference, 2014.
  18. 18.Gordon Christie, Neil Fendley, James Wilson, and Ryan Mukherjee. Functional map of the world. In Conference on Computer Vision and Pattern Recognition (CVPR), 2018. https://arxiv.org/abs/1711.07846.
  19. 19.Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Conference on Computer Vision and Pattern Recognition (CVPR), 2014. https://arxiv.org/abs/1311.3618.
  20. 20.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. RandAugment: Practical automated data augmentation with a reduced search space. In Conference on Computer Vision and Pattern Recognition (CVPR), 2020. https://arxiv.org/abs/1909.13719.
  21. 21.Ido Dagan, Oren Glickman, and Bernardo Magnini. The pascal recognising textual entailment challenge. In Machine Learning Challenges Workshop, 2005.
  22. 22.Zihang Dai, Hanxiao Liu, Quoc Le, and Mingxing Tan. CoAtNet: Marrying convolution and attention for all data sizes. Advances in Neural Information Processing Systems, 34, 2021.
  23. 23.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Conference on Computer Vision and Pattern Recognition, 2009. https://ieeexplore.ieee.org/document/5206848.
  24. 24.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics (NAACL), 2019a. URL https://aclanthology.org/N19-1423.
  25. 25.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019b. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  26. 26.Thomas G Dietterich. Ensemble methods in machine learning. In International workshop on multiple classifier systems, 2000. https://link.springer.com/chapter/10.1007/3-540-45014-9_1.
  27. 27.Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith. Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping. arXiv preprint arXiv:2002.06305, 2020.
  28. 28.Bill Dolan and Chris Brockett. Automatically constructing a corpus of sentential paraphrases. In Proc. of IWP, 2005.
  29. 29.Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell. Decaf: A deep convolutional activation feature for generic visual recognition. In International conference on machine learning, pages 647–655. PMLR, 2014.
  30. 30.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021. https://arxiv.org/abs/2010.11929.
  31. 31.Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization. In International Conference on Learning Representations, 2021. https://openreview.net/forum?id=6Tm1mposlrM.
  32. 32.Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning (ICML), 2020. https://arxiv.org/abs/1912.05671.
  33. 33.Yoav Freund and Robert E Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of Computer and System Sciences, 1997. https://www.sciencedirect.com/science/article/pii/S002200009791504X.
  34. 34.Jerome Friedman, Trevor Hastie, Robert Tibshirani, et al. The elements of statistical learning. Springer series in statistics New York, 2001.
  35. 35.Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. Clip-adapter: Better vision-language models with feature adapters. arXiv preprint arXiv:2110.04544, 2021.
  36. 36.Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry Vetrov, and Andrew Gordon Wilson. Loss surfaces, mode connectivity, and fast ensembling of dnns. In Advances in Neural Information Processing Systems (NeurIPS), 2018. https://arxiv.org/abs/1802.10026.
  37. 37.Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. The third pascal recognizing textual entailment challenge. In Proc. of the ACL-PASCAL workshop on textual entailment and paraphrasing, 2007.
  38. 38.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 580–587, 2014.
  39. 39.Raphael Gontijo-Lopes, Yann Dauphin, and Ekin Dogus Cubuk. No one representation to rule them all: Overlapping features of training methods. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=BK-4qbGgIE3.
  40. 40.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In International Conference on Machine Learning (ICML), 2017. https://arxiv.org/abs/1706.04599.
  41. 41.Yunhui Guo, Honghui Shi, Abhishek Kumar, Kristen Grauman, Tajana Rosing, and Rogerio Feris. Spottune: transfer learning through adaptive fine-tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4805–4814, 2019.
  42. 42.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. International Conference on Computer Vision (ICCV), 2021a. https://arxiv.org/abs/2006.16241.
  43. 43.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. Conference on Computer Vision and Pattern Recognition (CVPR), 2021b. https://arxiv.org/abs/1907.07174.
  44. 44.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Dark knowledge, 2014. https://www.ttic.edu/dl/dark14.pdf.
  45. 45.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. In Advances in Neural Information Processing Systems (NeurIPS) Deep Learning Workshop, 2015. https://arxiv.org/abs/1503.02531.
  46. 46.Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. Averaging weights leads to wider optima and better generalization. In Conference on Uncertainty in Artificial Intelligence (UAI), 2018. https://arxiv.org/abs/1803.05407.
  47. 47.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2102.05918.
  48. 48.Marcin Junczys-Dowmunt, Tomasz Dwojak, and Rico Sennrich. The amu-uedin submission to the wmt16 news translation task: Attention-based nmt models as feature functions in phrase-based smt. arXiv preprint arXiv:1605.04809, 2016.
  49. 49.Jean Kaddour, Linqing Liu, Ricardo Silva, and Matt J Kusner. Questions for flat-minima optimization of modern neural networks. arXiv preprint arXiv:2202.00661, 2022.
  50. 50.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2014. https://arxiv.org/abs/1412.6980.
  51. 51.Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A. Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang. WILDS: A benchmark of in-the-wild distribution shifts. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2012.07421.
  52. 52.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. In European Conference on Computer Vision (ECCV), 2020. https://arxiv.org/abs/1912.11370.
  53. 53.Simon Kornblith, Jonathon Shlens, and Quoc V Le. Do better imagenet models transfer better? In Conference on Computer Vision and Pattern Recognition (CVPR), 2019. https://arxiv.org/abs/1805.08974.
  54. 54.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In International Conference on Computer Vision (ICCV) Workshops, 2013. https://ieeexplore.ieee.org/document/6755945.
  55. 55.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images, 2009. https://www.cs.toronto.edu/˜kriz/learning-features-2009-TR.pdf.
  56. 56.Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=UYneFzXSJWh.
  57. 57.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In Advances in Neural Information Processing Systems (NeurIPS), 2017. https://arxiv.org/abs/1612.01474.
  58. 58.Yann LeCun. The mnist database of handwritten digits. http://yann.lecun.com/exdb/mnist/, 1998.
  59. 59.Julien-Charles Lévesque, Christian Gagné, and Robert Sabourin. Bayesian hyperparameter optimization for ensemble learning. arXiv preprint arXiv:1605.06394, 2016.
  60. 60.Xingjian Li, Haoyi Xiong, Haozhe An, Cheng-Zhong Xu, and Dejing Dou. Rifle: Backpropagation in depth for deep transfer learning through re-initializing the fully-connected layer. In International Conference on Machine Learning, pages 6010–6019. PMLR, 2020.
  61. 61.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations (ICLR), 2016. https://arxiv.org/abs/1608.03983.
  62. 62.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019. https://openreview.net/forum?id=Bkg6RiCqY7.
  63. 63.Edward Ma. Nlp augmentation. https://github.com/makcedward/nlpaug, 2019.
  64. 64.Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens Van Der Maaten. Exploring the limits of weakly supervised pretraining. In European Conference on Computer Vision (ECCV), 2018. https://arxiv.org/abs/1805.00932.
  65. 65.Michael Matena and Colin Raffel. Merging models with fisher-weighted averaging, 2021. https://arxiv.org/abs/2111.09832.
  66. 66.Brian W Matthews. Comparison of the predicted and observed secondary structure of t4 phage lysozyme. Biochimica et Biophysica Acta (BBA)-Protein Structure, 1975.
  67. 67.Hector Mendoza, Aaron Klein, Matthias Feurer, Jost Tobias Springenberg, and Frank Hutter. Towards automatically-tuned neural networks. In Workshop on Automatic Machine Learning, pages 58–65. PMLR, 2016.
  68. 68.Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie. Slip: Self-supervision meets language-image pre-training. arXiv preprint arXiv:2112.12750, 2021.
  69. 69.Basil Mustafa, Carlos Riquelme, Joan Puigcerver, André Susano Pinto, Daniel Keysers, and Neil Houlsby. Deep ensembles for low-data transfer learning, 2020. https://arxiv.org/abs/2010.06866.
  70. 70.Vaishnavh Nagarajan and J. Zico Kolter. Uniform convergence may be unable to explain generalization in deep learning. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/05e97c207235d63ceb1db43c60db7bbb-Paper.pdf.
  71. 71.Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang. What is being transferred in transfer learning? In Advances in Neural Information Processing Systems (NeurIPS), 2020. https://arxiv.org/abs/2008.11687.
  72. 72.Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua V Dillon, Balaji Lakshminarayanan, and Jasper Snoek. Can you trust your model's uncertainty? evaluating predictive uncertainty under dataset shift. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1906.02530.
  73. 73.Hieu Pham, Zihang Dai, Golnaz Ghiasi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, and Quoc V. Le. Combined scaling for zero-shot transfer learning, 2021. https://arxiv.org/abs/2111.10050.
  74. 74.Boris Teodorovich Polyak. New method of stochastic approximation type. Automation and remote control, 1990.
  75. 75.Ofir Press and Lior Wolf. Using the output embedding to improve language models. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 157–163, Valencia, Spain, April 2017. Association for Computational Linguistics. URL https://aclanthology.org/E17-2025.
  76. 76.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML), 2021. https://arxiv.org/abs/2103.00020.
  77. 77.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 2020a. http://jmlr.org/papers/v21/20-074.html.
  78. 78.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020b. URL http://jmlr.org/papers/v21/20-074.html.
  79. 79.Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do ImageNet classifiers generalize to ImageNet? In International Conference on Machine Learning (ICML), 2019. https://arxiv.org/abs/1902.10811.
  80. 80.Rebecca Roelofs, Nicholas Cain, Jonathon Shlens, and Michael C Mozer. Mitigating bias in calibration error estimation, 2020. https://arxiv.org/abs/2012.08668.
  81. 81.David Ruppert. Efficient estimations from a slowly convergent robbins-monro process, 1988. https://ecommons.cornell.edu/bitstream/handle/1813/8664/TR000781.pdf.
  82. 82.Tonmoy Saikia, Thomas Brox, and Cordelia Schmid. Optimized generic feature learning for few-shot classification across domains. arXiv preprint arXiv:2001.07926, 2020.
  83. 83.Vaishaal Shankar, Rebecca Roelofs, Horia Mania, Alex Fang, Benjamin Recht, and Ludwig Schmidt. Evaluating machine accuracy on imagenet. In International Conference on Machine Learning (ICML), 2020. http://proceedings.mlr.press/v119/shankar20c/shankar20c.pdf.
  84. 84.Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson. Cnn features off-the-shelf: an astounding baseline for recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2014. https://arxiv.org/abs/1403.6382.
  85. 85.Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR, 2018.
  86. 86.Yang Shu, Zhi Kou, Zhangjie Cao, Jianmin Wang, and Mingsheng Long. Zoo-tuning: Adaptive transfer from a zoo of models. In International Conference on Machine Learning, pages 9626–9637. PMLR, 2021.
  87. 87.Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Mostofa Patwary, Mr Prabhat, and Ryan Adams. Scalable bayesian optimization using deep neural networks. In International conference on machine learning, pages 2171–2180. PMLR, 2015.
  88. 88.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of EMNLP, 2013.
  89. 89.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016.
  90. 90.Tijmen Tieleman and Geoffrey Hinton. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning, 4(2):26–31, 2012.
  91. 91.Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Herve Jegou. Fixing the train-test resolution discrepancy. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://proceedings.neurips.cc/paper/2019/file/d03a857a23b5285736c4d55e0bb067c8-Paper.pdf.
  92. 92.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  93. 93.Johannes Von Oswald, Seijin Kobayashi, Joao Sacramento, Alexander Meulemans, Christian Henning, and Benjamin F Grewe. Neural networks with late-phase weights. arXiv preprint arXiv:2007.12927, 2020.
  94. 94.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
  95. 95.Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing. Learning robust global representations by penalizing local predictive power. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1905.13549.
  96. 96.Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. Neural network acceptability judgments. TACL, 7:625–641, 2019.
  97. 97.Jason Wei and Kai Zou. Eda: Easy data augmentation techniques for boosting performance on text classification tasks. arXiv preprint arXiv:1901.11196, 2019.
  98. 98.Florian Wenzel, Jasper Snoek, Dustin Tran, and Rodolphe Jenatton. Hyperparameter ensembles for robustness and uncertainty quantification. arXiv preprint arXiv:2006.13570, 2020.
  99. 99.Ross Wightman. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  100. 100.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online, October 2020. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/2020.emnlp-demos.6.
  101. 101.Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Simon Kornblith, Rebecca Roelofs, Raphael Gontijo-Lopes, Hannaneh Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt. Robust fine-tuning of zero-shot models. 2021. https://arxiv.org/abs/2109.01903.
  102. 102.Jianxiong Xiao, Krista A Ehinger, James Hays, Antonio Torralba, and Aude Oliva. Sun database: Exploring a large collection of scene categories. International Journal of Computer Vision, 2016. https://link.springer.com/article/10.1007/s11263-014-0748-y.
  103. 103.LI Xuhong, Yves Grandvalet, and Franck Davoine. Explicit inductive bias for transfer learning with convolutional networks. In International Conference on Machine Learning, pages 2825–2834. PMLR, 2018.
  104. 104.I Zeki Yalniz, Hervé Jégou, Kan Chen, Manohar Paluri, and Dhruv Mahajan. Billion-scale semi-supervised learning for image classification, 2019. https://arxiv.org/abs/1905.00546.
  105. 105.Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. How transferable are features in deep neural networks? In Advances in Neural Information Processing Systems (NeurIPS), 2014. https://arxiv.org/abs/1411.1792.
  106. 106.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022.
  107. 107.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6023–6032, 2019.
  108. 108.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers, 2021. https://arxiv.org/abs/2106.04560.
  109. 109.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. 2017. https://arxiv.org/abs/1710.09412.
  110. 110.Michael R Zhang, James Lucas, Geoffrey Hinton, and Jimmy Ba. Lookahead optimizer: k steps forward, 1 step back. In Advances in Neural Information Processing Systems (NeurIPS), 2019. https://arxiv.org/abs/1907.08610.
  111. 111.Renrui Zhang, Rongyao Fang, Peng Gao, Wei Zhang, Kunchang Li, Jifeng Dai, Yu Qiao, and Hongsheng Li. Tip-adapter: Training-free clip-adapter for better vision-language modeling. arXiv preprint arXiv:2111.03930, 2021.
  112. 112.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models, 2021. https://arxiv.org/abs/2109.01134.

Citation

MLA
Wortsman, M., et al. “Model Soups: Averaging Weights of Multiple Fine-tuned Models Improves Accuracy Without Increasing Inference Time”. arXiv, 2022, http://arxiv.org/abs/2203.05482v3.
APA
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., & Schmidt, L. (2022). Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. arXiv. http://arxiv.org/abs/2203.05482v3
Chicago
Wortsman, M., G. Ilharco, S. Y. Gadre, et al. 2022. “Model Soups: Averaging Weights of Multiple Fine-tuned Models Improves Accuracy Without Increasing Inference Time”. arXiv. http://arxiv.org/abs/2203.05482v3.
Harvard
Wortsman, M. et al. (2022) “Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.05482v3.
Vancouver
1. Wortsman M, Ilharco G, Gadre SY, et al (2022) Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. arXiv

BibTeX

@article{wortsman2022model,
  title = {Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time},
  author = {Wortsman, Mitchell and Ilharco, Gabriel and Gadre, Samir Yitzhak and Roelofs, Rebecca and Gontijo-Lopes, Raphael and Morcos, Ari S. and Namkoong, Hongseok and Farhadi, Ali and Carmon, Yair and Kornblith, Simon and Schmidt, Ludwig},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.05482v3},
  eprint = {2203.05482}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/