AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning

Yaqing WangSahaj AgarwalSubhabrata MukherjeeXiaodong LiuJing GaoAhmed Hassan AwadallahJianfeng Gao

article2022EMNLP210 citationsDistinguished Paper Award

Proposes AdaMix, a general parameter-efficient fine-tuning method that uses stochastic routing and module merging across multiple adaptation components to outperform full model fine-tuning on natural language understanding and generation benchmarks while tuning only 0.1 to 0.2 percent of parameters.

Listen

Adapting large pre-trained language models for specialized business tasks usually requires full fine-tuning, which updates hundreds of millions to billions of parameters. This creates massive storage, memory, and operational costs because an entire copy of the model must be saved and served for each new application. To mitigate this expense, parameter-efficient fine-tuning (PEFT) methods update only tiny add-on components while keeping the main model frozen. However, conventional parameter-efficient approaches historically lag behind the accuracy of full-model tuning.

The article introduces and evaluates AdaMix, a general framework designed to improve parameter-efficient fine-tuning without increasing deployment computational costs or memory requirements. The primary objective is to demonstrate that introducing a mixture of multiple lightweight adaptation modules during training can match or exceed the performance of full fine-tuning while adjusting only 0.1% to 0.2% of the total model parameters.

The authors conducted extensive empirical evaluations across eight natural language understanding tasks from the standard GLUE benchmark, three text-generation datasets, and multiple few-shot learning scenarios with minimal labeled data. They tested AdaMix on widely used base models, including BERT, RoBERTa, and GPT-2, comparing it directly against full-model tuning and leading parameter-efficient baselines. The AdaMix approach trains multiple adaptation sub-modules by randomly routing inputs through them, enforces stability via consistency regularization, shares certain parameters, and ultimately merges all modules into a single standard module through weight averaging at inference.

The key findings show substantial performance and efficiency gains. First, AdaMix achieved an average score of 89.9 on GLUE using a RoBERTa-large encoder with only 0.8 million tuned parameters, outperforming both full-model fine-tuning (88.9 with 355 million parameters) and baseline adapters (86.4 to 88.6). Second, it proved effective across text generation tasks, improving generation quality metrics over strong baselines like LoRA and standard adapters on datasets such as E2E, DART, and WebNLG. Third, in few-shot settings with only 30 labeled training examples, AdaMix attained an average score of 79.3, exceeding full prompt fine-tuning (77.5) and existing efficient adapters (57.8 to 77.6). Fourth, weight merging during deployment outperformed costly multi-pass ensembling while keeping operational computation identical to standard single-adapter setups.

These results demonstrate that organizations can reduce task-specific storage overhead by up to 444 times (lowering per-task storage from 355 megabytes to 0.8 megabytes) without compromising model accuracy. By decoupling training-time multi-view learning from inference-time simplicity, AdaMix lowers serving latency and cloud memory costs, making enterprise-scale deployment of specialized language models significantly more economical.

Decision-makers should consider adopting the AdaMix architecture when deploying multi-task language systems, particularly where hosting separate full-scale models is cost-prohibitive. As a practical next step, technical teams can pilot AdaMix on existing adapter- or low-rank-based pipelines to validate task performance before wide-scale deployment. However, leaders should note that AdaMix requires approximately 1 to 2 times more training iterations than standard parameter-efficient methods, resulting in higher upfront compute and carbon costs during the training phase. Overall confidence in the evaluation is high across the tested architectures, though further validation is recommended before extending the technique to newer parameter-efficient mechanisms like prefix- or prompt-tuning.

Cover for AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning

Abstract

Standard fine-tuning of large pre-trained language models (PLMs) for downstream tasks requires updating hundreds of millions to billions of parameters, and storing a large copy of the PLM weights for every task resulting in increased cost for storing, sharing and serving the models. To address this, parameter-efficient fine-tuning (PEFT) techniques were introduced where small trainable components are injected in the PLM and updated during fine-tuning. We propose AdaMix as a general PEFT method that tunes a mixture of adaptation modules – given the underlying PEFT method of choice – introduced in each Transformer layer while keeping most of the PLM weights frozen. For instance, AdaMix can leverage a mixture of adapters like Houlsby (Houlsby et al., 2019) or a mixture of low rank decomposition matrices like LoRA (Hu et al., 2021) to improve downstream task performance over the corresponding PEFT methods for fully supervised and few-shot NLU and NLG tasks. Further, we design AdaMix such that it matches the same computational cost and the number of tunable parameters as the underlying PEFT method. By only tuning 0.1 – 0.2% of PLM parameters, we show that AdaMix outperforms SOTA parameter-efficient fine-tuning and full model fine-tuning for both NLU and NLG tasks. Code and models are made available at https://aka.ms/AdaMix.

Table of Contents

  • 1 Introduction
  • 2 Mixture-of-Adaptations
  • 2.1 Routing Policy
  • 2.2 Consistency regularization
  • 2.3 Adaptation module merging
  • 2.4 Adaptation module sharing
  • 2.5 Connection to Bayesian Neural Networks and Model Ensembling
  • 3 Experiments
  • 3.1 Experimental Setup
  • 3.2 Key Results
  • 3.2.1 NLU Tasks
  • 3.2.2 NLG Tasks
  • 3.2.3 Few-shot NLU
  • 3.3 Ablation Study
  • 4 Related Work
  • 5 Conclusions
  • 6 Limitations
  • 7 Acknowledgment
  • References
  • Appendix
  • A Background
  • A.1 Mixture-of-Experts
  • A.2 Adapters
  • B Few-shot NLU Datasets
  • C Ablation Study
  • D Detailed Results on NLU Tasks
  • E Hyper-parameter

Knowls

  1. Knowl 1 — AdaMix mixture-of-adaptations framework

    model/method

    AdaMix is a parameter-efficient fine-tuning framework that replaces the single adaptation module normally inserted into each of the LL Transformer layers with a mixture of MM adaptation modules, while keeping the pretrained language-model parameters frozen. For adapter-based tuning, module AijA_{ij} in layer ii contains a down-projection Wijdown∈Rd×rW^{\mathrm{down}}_{ij}\in\mathbb{R}^{d\times r} and an up-projection Wijup∈Rr×dW^{\mathrm{up}}_{ij}\in\mathbb{R}^{r\times d}, where dd is the Transformer hidden size and rr is the adapter bottleneck dimension. For a hidden representation x∈Rdx\in\mathbb{R}^{d}, the selected adapter contributes a residual transformation of the form

    x′=x+f ⁣(xWdown)Wup,x' = x + f\!\left(xW^{\mathrm{down}}\right)W^{\mathrm{up}},

    where ff is the adapter nonlinearity. The same mixture construction can be applied to other PEFT mechanisms; for LoRA, AdaMix maintains multiple low-rank adaptation matrices and routes inputs among them. AdaMix trains the mixture as multiple task views, but uses one selected adaptation path at a time and merges the learned modules before inference, so the deployed model has the computational cost and adaptation-parameter count of a single underlying PEFT module.

  2. Knowl 2 — Stochastic routing and adaptation-module sharing

    model/method

    During AdaMix training, a pair of adaptation projections is selected randomly in every Transformer layer at each training step. The selected pair can contain one up-projection and one down-projection from the mixture; every example in the current input batch uses the same selected modules. Random routing requires no learned gating network, adds no routing parameters or routing computation, and activates only one adaptation path, so the per-example FLOPs match those of a conventional single adapter or single LoRA path. The randomly selected modules consequently learn different transformations of the same downstream task.

    To make this multi-view training more effective for low-resource tasks, AdaMix shares part of the adaptation modules. In the standard adapter configuration, all mixture components share the up-projection, Wijup=WiupW^{\mathrm{up}}_{ij}=W^{\mathrm{up}}_i, while the down-projections remain distinct. This reduces the number of independently learned parameters and improves training efficiency and convergence when labeled data are scarce.

  3. Knowl 3 — Consistency regularization across stochastic routes

    equation

    AdaMix trains two independently routed versions of the same input so that different adaptation-module choices produce consistent predictions. Let xx be the representation produced by the frozen pretrained language model, let AA and BB denote the two sets of adaptation modules selected across the LL Transformer layers, let zA(x)z^A(x) and zB(x)z^B(x) be their predicted class-logit vectors, let CC be the number of classes, and let I(x,c)∈{0,1}I(x,c)\in\{0,1\} indicate whether class cc is the correct label. The paper defines the training objective as

    L=−(∑c=1CI(x,c)log⁡softmax⁡ ⁣(zcA(x))+12[KL⁡ ⁣(z(.)A(x) ∥ z(.)B(x))+KL⁡ ⁣(z(.)B(x) ∥ z(.)A(x))]),\mathcal{L}=-\left(\sum_{c=1}^{C} I(x,c)\log\operatorname{softmax}\!\left(z^A_c(x)\right) + \frac{1}{2}\left[\operatorname{KL}\!\left(z^A_{(.)}(x)\,\|\,z^B_{(.)}(x)\right)+\operatorname{KL}\!\left(z^B_{(.)}(x)\,\|\,z^A_{(.)}(x)\right)\right]\right),

    where KL⁡\operatorname{KL} is the Kullback–Leibler divergence between the two routed predictions. Only adaptation parameters are updated; the pretrained language-model parameters remain frozen. The symmetric consistency term prevents the stochastic modules from diverging into incompatible predictors and reduces the instability caused by random route selection.

  4. Knowl 4 — Inference-time adaptation-module merging

    model/method

    AdaMix collapses the separately trained adaptation modules into one module per Transformer layer during inference. For layer ii with MM up-projection matrices WijupW^{\mathrm{up}}_{ij} and MM down-projection matrices WijdownW^{\mathrm{down}}_{ij}, where j∈{1,…,M}j\in\{1,\ldots,M\} indexes mixture components, the merged projections are computed by elementwise weight averaging:

    Wi′up=1M∑j=1MWijup,Wi′down=1M∑j=1MWijdown.W^{\prime\mathrm{up}}_i=\frac{1}{M}\sum_{j=1}^{M}W^{\mathrm{up}}_{ij},\qquad W^{\prime\mathrm{down}}_i=\frac{1}{M}\sum_{j=1}^{M}W^{\mathrm{down}}_{ij}.

    The merged projections replace the mixture at serving time, so inference stores and evaluates only one adaptation module in each layer. This removes the extra serving and storage cost of retaining all stochastic modules while preserving the performance benefit learned from their multiple task views. The same averaging principle is applied to the corresponding low-rank matrices when LoRA is the underlying PEFT mechanism.

  5. Knowl 5 — Connection to Bayesian averaging and model ensembling

    model/method

    AdaMix interprets stochastic routing as a tractable approximation to Bayesian marginalization over adaptation weights. Let fW(x)f_W(x) be the output of a model with adaptation weights WW, let qθ(W)q_\theta(W) be the stochastic-routing distribution, and let W~t∼qθ(W)\widetilde W_t\sim q_\theta(W) be the masked adaptation weights sampled on pass tt. For a classification label cc, the Monte Carlo approximation used in the paper is

    p(y=c∣x)≈1T∑t=1Tp ⁣(y=c∣fW~t(x))=1T∑t=1Tsoftmax⁡ ⁣(fW~t(x))c,p(y=c\mid x)\approx\frac{1}{T}\sum_{t=1}^{T}p\!\left(y=c\mid f_{\widetilde W_t}(x)\right)=\frac{1}{T}\sum_{t=1}^{T}\operatorname{softmax}\!\left(f_{\widetilde W_t}(x)\right)_c,

    where TT is the number of stochastic routing passes and yy is the output label. Direct stochastic or logit-level ensembling requires retaining and evaluating all TT routed models. AdaMix instead averages the adaptation weights, approximately replacing the ensemble by a single model with W~=1T∑tW~t\widetilde W=\frac{1}{T}\sum_t\widetilde W_t. The paper argues, based on the relationship between weight averaging and logit ensembling, that this can retain favorable performance when the solutions lie in a sufficiently flat region and predictions are confident, while requiring only one merged copy at inference.

  6. Knowl 6 — Supervised NLU performance on GLUE

    empirical result

    The authors evaluate AdaMix on the eight GLUE development tasks using RoBERTa-large and BERT-base encoders. The main adapter configuration uses four mixture modules, bottleneck dimension 1616 for RoBERTa-large and 4848 for BERT-base. On RoBERTa-large, AdaMix uses only 0.80.8M tunable parameters, approximately 0.23%0.23\% of the 355.0355.0M-parameter encoder, yet obtains the highest reported average score, 89.989.9, compared with 88.988.9 for full fine-tuning and 88.688.6 for LoRA. The task-level comparison is:

    Model #Param. MNLI QNLI SST-2 QQP MRPC CoLA RTE STS-B Avg.
    Full fine-tuning 355.0M 90.2 94.7 96.4 92.2 90.9 68.0 86.6 92.4 88.9
    Pfeiffer, 3.0M 3.0M 90.2 94.8 96.1 91.9 90.2 68.3 83.8 92.1 88.4
    Pfeiffer, 0.8M 0.8M 90.5 94.8 96.6 91.7 89.7 67.8 80.1 91.9 87.9
    Houlsby, 6.0M 6.0M 89.9 94.7 96.2 92.1 88.7 66.5 83.4 91.0 87.8
    Houlsby, 0.8M 0.8M 90.3 94.7 96.3 91.5 87.7 66.3 72.9 91.5 86.4
    LoRA 0.8M 90.6 94.8 96.2 91.6 90.2 68.2 85.2 92.3 88.6
    AdaMix Adapter 0.8M 90.9 95.4 97.1 92.3 91.9 70.2 89.2 92.4 89.9

    With BERT-base, AdaMix reaches an average of 84.584.5 using 0.90.9M tunable parameters, compared with 82.782.7 for full fine-tuning and 83.583.5 for the strongest listed UNIPELT baseline. The paper reports AdaMix as the only PEFT method that surpasses full fine-tuning on the complete GLUE evaluation and average score in its RoBERTa-large comparison.

  7. Knowl 7 — Natural-language-generation performance with adapters and LoRA

    empirical result

    AdaMix was evaluated with GPT-2 medium on the E2E, DART, and WebNLG generation tasks, using both adapters and LoRA as the underlying PEFT mechanism. The NLG configurations use eight adaptation modules; AdaMix tunes only 0.350.35–0.420.42M parameters, approximately 0.1%0.1\% of the 354.92354.92M-parameter backbone. AdaMix LoRA obtains the best listed result on all three datasets, and AdaMix Adapter also improves over the corresponding conventional adapter.

    E2E
    Model #Param. BLEU NIST METEOR ROUGE-L CIDEr
    Full fine-tuning 354.92M 68.2 8.62 46.2 71.0 2.47
    Houlsby Adapter 11.09M 67.3 8.50 46.0 70.7 2.44
    LoRA 0.35M 70.4 8.85 46.8 71.8 2.53
    AdaMix Adapter 0.42M 69.8 8.75 46.8 71.9 2.52
    AdaMix LoRA 0.35M 71.0 8.89 46.8 72.2 2.54
    Dataset Model #Param. BLEU
    DART Full fine-tuning 354.92M 46.2
    DART LoRA 0.35M 47.35
    DART AdaMix Adapter 0.42M 47.72
    DART AdaMix LoRA 0.35M 47.86
    WebNLG Full fine-tuning 354.92M 46.5
    WebNLG LoRA 0.35M 55.37
    WebNLG AdaMix Adapter 0.42M 54.94
    WebNLG AdaMix LoRA 0.35M 55.64

    These results demonstrate that the mixture framework improves both adapter-based and low-rank adaptation rather than being tied to one particular PEFT parameterization.

  8. Knowl 8 — Few-shot NLU results with 30 labeled examples

    empirical result

    For few-shot NLU, the authors use RoBERTa-large as a frozen encoder, prompt-based fine-tuning, and ∣K∣=30|K|=30 labeled examples per task on six GLUE-related tasks: MNLI, RTE, QQP, SST-2, Subj, and MPQA. No validation set or additional unlabeled data is used; results are averaged over five random seeds. AdaMix Adapter obtains an average score of 79.379.3, exceeding full prompt fine-tuning at 77.577.5 and LiST Adapter at 77.677.6, while tuning the same reported 1414M adapter parameters as those adapter baselines.

    Model MNLI RTE QQP SST-2 Subj MPQA Avg.
    Full Prompt Fine-tuning 62.8(2.6) 66.1(2.2) 71.1(1.5) 91.5(1.0) 91.0(0.5) 82.7(3.8) 77.5
    Head-only 54.1(1.1) 58.8(2.6) 56.7(4.5) 85.6(1.0) 82.1(2.5) 64.1(2.1) 66.9
    BitFit 54.4(1.3) 59.8(3.5) 58.6(4.4) 87.3(1.1) 83.9(2.3) 65.8(1.8) 68.3
    Prompt-tuning 47.3(0.2) 53.0(0.6) 39.9(0.7) 75.7(1.7) 51.5(1.4) 70.9(2.4) 56.4
    Houlsby Adapter 35.7(1.1) 51.0(3.0) 62.8(3.0) 57.0(6.2) 83.2(5.4) 57.2(3.5) 57.8
    LiST Adapter 62.4(1.7) 66.6(3.9) 71.2(2.6) 91.7(1.0) 90.9(1.3) 82.6(2.0) 77.6
    AdaMix Adapter 65.6(2.6) 69.6(3.4) 72.6(1.2) 91.8(1.1) 91.5(2.0) 84.7(1.6) 79.3

    The gains over full prompt fine-tuning are consistent with the fully supervised GLUE results, although the authors note that AdaMix and LiST use similar adapter designs in this setting.

  9. Knowl 9 — Ablation findings on merging, consistency, sharing, and mixture size

    empirical result

    Ablations on GLUE show that each component of AdaMix contributes to its performance. With a BERT-base encoder, the complete method with merging reaches an average score of 84.584.5, whereas removing merging and using random routing, fixed routing, or four-pass logit ensembling gives 83.383.3, 83.783.7, and 83.283.2, respectively. Weight merging therefore performs better than retaining random routes or ensembling their logits, while requiring only 0.90.9M inference parameters instead of 3.63.6M for the four-route alternatives.

    Variant #Param. Avg. GLUE score
    Full fine-tuning 110M 82.7
    AdaMix with merging 0.9M 84.5
    AdaMix without merging + random routing 3.6M 83.3
    AdaMix without merging + fixed routing 0.9M 83.7
    AdaMix without merging + logit ensemble 3.6M 83.2

    Removing consistency regularization lowers performance on the low-resource RTE task from 89.289.2 to 84.884.8, and removing sharing lowers it to 84.184.1. Across MNLI, QNLI, SST-2, MRPC, and RTE, the complete method scores 90.990.9, 95.495.4, 97.197.1, 91.991.9, and 89.289.2; the no-consistency variant scores 90.790.7, 95.095.0, 97.197.1, 91.491.4, and 84.884.8; and the no-sharing variant scores 90.990.9, 95.095.0, 96.496.4, 90.490.4, and 84.184.1. The performance gap caused by removing sharing becomes larger as the dataset size decreases.

    Increasing the number of modules from 22 to 44 improves or preserves performance on the tested RoBERTa-large tasks, but increasing it to 88 gives diminishing returns and hurts lower-resource tasks. For example, the MNLI/QNLI/SST-2/MRPC/RTE scores are 90.9/95.2/96.8/90.9/87.490.9/95.2/96.8/90.9/87.4 with 22 modules, 90.9/95.4/97.1/91.9/89.290.9/95.4/97.1/91.9/89.2 with 44, and 90.9/95.3/96.9/91.4/87.490.9/95.3/96.9/91.4/87.4 with 88. Increasing adapter bottleneck dimension likewise helps initially but shows diminishing returns: with RoBERTa-large, dimensions 88, 1616, and 3232 use 0.40.4M, 0.80.8M, and 1.51.5M parameters and obtain average task behavior of 90.7/95.2/96.8/91.2/87.790.7/95.2/96.8/91.2/87.7, 90.9/95.4/97.1/91.9/89.290.9/95.4/97.1/91.9/89.2, and 91.0/95.4/96.8/90.7/89.291.0/95.4/96.8/90.7/89.2 on MNLI/QNLI/SST-2/MRPC/RTE.

  10. Knowl 10 — Training cost and untested PEFT combinations

    limitation

    AdaMix increases training cost relative to standard single-module PEFT because multiple adaptation-module copies are maintained during training. The authors report that AdaMix generally requires between 11 and 22 times as many training iterations as standard PEFT, creating an additional computational and carbon-footprint burden even though merged inference has single-module cost. The experiments cover adapters and LoRA as underlying PEFT mechanisms, but do not test AdaMix with prompt-tuning, prefix-tuning, or other PEFT combinations; the claimed generality to those methods remains future work.

Coverage note — Background, related work, acknowledgements, and proof-free implementation details not needed to reconstruct the contribution were omitted; the main appendix tables were represented through the load-bearing ablations and results rather than reproduced in full.

References

  1. 1.Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer. 2021. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7319–7328, Online. Association for Computational Linguistics.
  2. 2.Roy Bar Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006. The second PASCAL recognising textual entailment challenge.
  3. 3.Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009. The fifth PASCAL recognizing textual entailment challenge. In TAC.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005. The PASCAL recognising textual entailment challenge. In the First International Conference on Machine Learning Challenges: Evaluating Predictive Uncertainty Visual Object Classification, and Recognizing Textual Entailment.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Volume 1 (Long and Short Papers), pages 4171–4186.
  7. 7.William Fedus, Barret Zoph, and Noam Shazeer. 2021. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. arXiv preprint arXiv:2101.03961.
  8. 8.Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. 2020. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning, pages 3259–3269. PMLR.
  9. 9.Yarin Gal and Zoubin Ghahramani. 2015. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. CoRR, abs/1506.02142.
  10. 10.Yarin Gal, Riashat Islam, and Zoubin Ghahramani. 2017. Deep Bayesian active learning with image data. In Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, pages 1183–1192. PMLR.
  11. 11.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Association for Computational Linguistics (ACL).
  12. 12.Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017. The webnlg challenge: Generating text from rdf data. In Proceedings of the 10th International Conference on Natural Language Generation, pages 124–133.
  13. 13.Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007. The third PASCAL recognizing textual entailment challenge. In the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing.
  14. 14.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pages 2790–2799. PMLR.
  15. 15.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  16. 16.Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. 2018. Averaging weights leads to wider optima and better generalization. arXiv preprint arXiv:1803.05407.
  17. 17.Jaejun Lee, Raphael Tang, and Jimmy Lin. 2019. What would elsa do? freezing layers during transformer fine-tuning. arXiv preprint arXiv:1911.03090.
  18. 18.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668.
  19. 19.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. CoRR, abs/2104.08691.
  20. 20.Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer. 2021. Base layers: Simplifying training of large, sparse models. In ICML.
  21. 21.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. CoRR, abs/2101.00190.
  22. 22.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  23. 23.Yuning Mao, Lambert Mathias, Rui Hou, Amjad Almahairi, Hao Ma, Jiawei Han, Wen-tau Yih, and Madian Khabsa. 2021. Unipelt: A unified framework for parameter-efficient language model tuning. arXiv preprint arXiv:2110.07577.
  24. 24.Michael Matena and Colin Raffel. 2021. Merging models with fisher-weighted averaging. arXiv preprint arXiv:2111.09832.
  25. 25.Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, et al. 2020. Dart: Open-domain structured data record to text generation. arXiv preprint arXiv:2007.02871.
  26. 26.Behnam Neyshabur, Hanie Sedghi, and Chiyuan Zhang. 2020. What is being transferred in transfer learning? Advances in neural information processing systems, 33:512–523.
  27. 27.Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2017. The e2e dataset: New challenges for end-to-end generation. arXiv preprint arXiv:1706.09254.
  28. 28.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts.
  29. 29.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. arXiv preprint arXiv:2105.11447.
  30. 30.Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021. Adapterfusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 487–503.
  31. 31.Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulic, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020. Adapterhub: A framework for adapting transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020): Systems Demonstrations, pages 46–54, Online. Association for Computational Linguistics.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683.
  33. 33.Stephen Roller, Sainbayar Sukhbaatar, Arthur D. Szlam, and Jason Weston. 2021. Hash layers for large sparse models. ArXiv, abs/2106.04426.
  34. 34.Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Raluca Turc, Jacob Eisenstein, Dipanjan Das, and Ellie Pavlick. 2022. The multiBERTs: BERT reproductions for robustness analysis. In International Conference on Learning Representations.
  35. 35.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538.
  36. 36.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990.
  37. 37.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. Recursive deep models for semantic compositionality over a sentiment treebank.
  38. 38.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826.
  39. 39.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  40. 40.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019. GLUE: A multi-task benchmark and analysis platform for natural language understanding.
  41. 41.Yaqing Wang, Subhabrata Mukherjee, Xiaodong Liu, Jing Gao, Ahmed Hassan Awadallah, and Jianfeng Gao. 2021. List: Lite self-training makes efficient few-shot learners. arXiv preprint arXiv:2110.06274.
  42. 42.Janyce Wiebe, Theresa Wilson, and Claire Cardie. 2005. Annotating expressions of opinions and emotions in language. Language resources and evaluation, 39(2):165–210.
  43. 43.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference.
  44. 44.Mitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs, Raphael Gontijo-Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al. 2022. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. arXiv preprint arXiv:2203.05482.
  45. 45.An Yang, Junyang Lin, Rui Men, Chang Zhou, Le Jiang, Xianyan Jia, Ang Wang, Jie Zhang, Jiamang Wang, Yong Li, et al. 2021. M6-t: Exploring sparse expert models and beyond. arXiv preprint arXiv:2105.15082.
  46. 46.Elad Ben Zaken, Shauli Ravfogel, and Yoav Goldberg. 2021. Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language models. arXiv preprint arXiv:2106.10199.
  47. 47.Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q Weinberger, and Yoav Artzi. 2021. Revisiting few-sample BERT fine-tuning.
  48. 48.Simiao Zuo, Xiaodong Liu, Jian Jiao, Young Jin Kim, Hany Hassan, Ruofei Zhang, Tuo Zhao, and Jianfeng Gao. 2021. Taming sparsely activated transformer with stochastic experts. arXiv preprint arXiv:2110.04260.

Citation

MLA
Wang, Y., et al. “AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning”. arXiv, 2022, http://arxiv.org/abs/2205.12410v2.
APA
Wang, Y., Agarwal, S., Mukherjee, S., Liu, X., Gao, J., Awadallah, A. H., & Gao, J. (2022). AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning. arXiv. http://arxiv.org/abs/2205.12410v2
Chicago
Wang, Y., S. Agarwal, S. Mukherjee, et al. 2022. “AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning”. arXiv. http://arxiv.org/abs/2205.12410v2.
Harvard
Wang, Y. et al. (2022) “AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.12410v2.
Vancouver
1. Wang Y, Agarwal S, Mukherjee S, Liu X, Gao J, Awadallah AH, Gao J (2022) AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning. arXiv

BibTeX

@article{wang2022adamix,
  title = {AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning},
  author = {Wang, Yaqing and Agarwal, Sahaj and Mukherjee, Subhabrata and Liu, Xiaodong and Gao, Jing and Awadallah, Ahmed Hassan and Gao, Jianfeng},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.12410v2},
  eprint = {2205.12410}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/