BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting

Zheng Xin YongHailey SchoelkopfNiklas MuennighoffAlham Fikri AjiDavid Ifeoluwa AdelaniKhalid AlmubarakM. Saiful BariLintang SutawikaJungo KasaiAhmed Baruwa

article2023ACL114 citations

Demonstrates that parameter-efficient adapter finetuning outperforms continued pretraining when adapting large multilingual models like BLOOM to unseen languages in resource-constrained, zero-shot prompting settings.

Listen

Large multilingual language models offer powerful capabilities, but training them from scratch requires immense computational resources. As a result, even leading models cover only a fraction of the world’s languages, excluding major populations and low-resource language communities. Because retraining entire models to add new languages is financially prohibitive, organizations need efficient techniques to expand language coverage without incurring unsustainable computing costs.

The article systematically evaluates different lightweight strategies to adapt the open-access BLOOM model family to eight previously unsupported languages. Its primary goal is to determine the most effective and computationally efficient methods for enabling zero-shot task performance in new languages under resource-constrained training budgets.

The authors conducted experimental benchmarks adapting BLOOM models ranging in scale from 560 million to 7.1 billion parameters across eight languages representing diverse scripts and families: Bulgarian, German, Greek, Guarani, Korean, Russian, Thai, and Turkish. The evaluation compared continued causal language pretraining against parameter-efficient modular approaches, primarily bottleneck adapters (MAD-X) and activation-scaling adapters ((IA)3). Training utilized a low-resource budget capped at 100,000 text samples per language (about 100 million tokens), without altering the base model’s subword vocabulary. Performance was measured across standard natural language understanding tasks, including inference, reasoning, and paraphrase detection, using zero-shot prompting.

The investigation produced several key findings regarding model scaling, architectural efficiency, and data requirements. First, adapter-based adaptation consistently outperforms continued pretraining for models with 3 billion or more parameters, while requiring substantially lower GPU memory and training runtime. Continued pretraining proved superior only for the smallest 560-million-parameter model. Second, the effectiveness of zero-shot prompting depends heavily on the volume of adaptation data rather than linguistic characteristics; unseen writing systems and diverse word orders adapted equally well, provided sufficient text was available. Third, effective adaptation requires approximately 100 million tokens; performance collapsed when training on 10,000 or fewer samples, explaining the poor results observed on the low-resource language Guarani (30,000 samples). Fourth, adapter capacity matters, as allocating more parameters to adapter modules steadily boosted task accuracy. Finally, for instruction-tuned model variants (BLOOMZ), directly incorporating target-language task data into the broader multitask training mixture proved far more effective than adapting the model solely on raw, unstructured monolingual text.

These results demonstrate that expanding large language models into new languages is viable without the extreme expense of full-model retraining. Relying on modular adapters significantly lowers compute infrastructure costs, shortens deployment timelines, and avoids catastrophic forgetting of original languages. This provides a practical path for enterprises and public institutions to democratize language technology across underrepresented languages.

Decision-makers seeking to extend large language models should deploy modular adapter frameworks rather than full continued pretraining when working with models of 3 billion parameters or larger. When deploying instruction-following assistants, teams should integrate target-language examples directly into diverse multitask fine-tuning mixtures. Projects should ensure a baseline training corpus of roughly 100 million tokens per language before attempting adaptation. If only small data volumes are available, organizations should run pilot benchmarks first, as low-data regimes can degrade baseline performance.

Confidence in these findings is high for classification and reasoning tasks across small- to medium-sized model tiers. However, caution is warranted when extrapolating these conclusions to text-generation tasks or extremely low-resource languages lacking sufficient text corpora. Additionally, computational constraints prevented testing on the maximum 176-billion-parameter BLOOM architecture, which remains an area for future validation.

Cover for BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting

Abstract

The BLOOM model is a large publicly available multilingual language model, but its pretraining was limited to 46 languages. To extend the benefits of BLOOM to other languages without incurring prohibitively large costs, it is desirable to adapt BLOOM to new languages not seen during pretraining. In this work, we apply existing language adaptation strategies to BLOOM and benchmark its zero-shot prompting performance on eight new languages in a resource-constrained setting. We find language adaptation to be effective at improving zero-shot performance in new languages. Surprisingly, we find that adapter-based finetuning is more effective than continued pretraining for large models. In addition, we discover that prompting performance is not significantly affected by language specifics, such as the writing system. It is primarily determined by the size of the language adaptation data. We also add new languages to BLOOMZ, which is a multitask finetuned version of BLOOM capable of following task instructions zero-shot. We find including a new language in the multitask fine-tuning mixture to be the most effective method to teach BLOOMZ a new language. We conclude that with sufficient training data language adaptation can generalize well to diverse languages. Our code is available at https://github.com/bigscience-workshop/multilingual-modeling.

Table of Contents

  • 1 Introduction
  • 1.1 Our Contributions
  • 2 Related Work
  • 3 Experimental settings
  • 3.1 BLOOM pretrained models
  • 3.2 New Languages
  • 3.3 Language Adaptation Strategies
  • 3.4 Language Adaptation Setting
  • 3.5 Tasks and Prompt Templates
  • 3.6 Baselines
  • 4 Results and Discussion
  • 4.1 Zero-shot Prompting Performance
  • 4.2 Perplexity
  • 4.3 Connection to Language Independent Representation
  • 4.4 Amount of Language Adaptation Data
  • 4.5 Adapters' Capacity
  • 4.6 Adapting BLOOMZ
  • 4.6.1 Adding Language Support through Unlabeled Data
  • 4.6.2 Adding Language Support through Instruction Tuning
  • 5 Conclusion
  • 6 Limitations
  • 6.1 Vocabulary and Embedding Adaptation
  • 6.2 Parameter-Efficient Finetuning Strategies
  • 6.3 Low-Resource Languages
  • 6.4 Generative Tasks
  • 6.5 Experimental Settings
  • References
  • Appendix
  • A Authors' Contributions
  • B How does Language Independent Representation changes with Model Sizes
  • C Batch Sizes
  • D Composable Sparse-Finetuning
  • E Korean PAWS-X
  • F Prompt Templates
  • G Other Parameter-Efficient Finetuning Strategies
  • H Language Adaptation Experimental Setup Details
  • I Number of Tokens for Language Adaptation Data
  • J Placement of Adapters
  • K Ablations
  • L Catastrophic Forgetting
  • M Pretraining Languages Existing in BLOOM
  • N Post-Hoc Experiments
  • O Artifacts

Knowls

  1. Knowl 1 — The best adaptation strategy depends on BLOOM model scale

    empirical result

    On zero-shot prompting for the tested new-language NLU tasks, continued pretraining gave the strongest results for BLOOM-560M, whereas adapter-based adaptation outperformed continued pretraining for models larger than 3 billion parameters. Among the adapter approaches, MAD-X produced higher average prompting performance, while (IA)³ used substantially fewer trainable parameters and trained faster on larger models. Thus, the study recommends continued pretraining for the smallest tested BLOOM model and adapters for larger tested models; the compute comparison used a single A100 GPU. When adaptation was instead applied to randomly initialized BLOOM, Russian XNLI performance remained near that of a random classifier, indicating that adaptation gains depended on the pretrained model.

  2. Knowl 2 — Evaluation setting for extending BLOOM to eight languages

    experimental setup

    The experiments adapted BLOOM models with 560 million, 1.1 billion, 1.7 billion, 3 billion, or 7.1 billion parameters to German, Bulgarian, Russian, Greek, Turkish, Thai, Korean, and Guarani. For each language, adaptation used at most 100,000 monolingual samples: deduplicated OSCAR data for seven languages and 30,000 Guarani sentences from the Jojajovai corpus. Training ran for 25,000 steps with batch size 8 and sequence length 1,024, corresponding to about 204 million training tokens. The byte-level BPE tokenizer was left unchanged. Evaluation was zero-shot prompting without task-specific finetuning, using translated XGLM prompt templates on NLU tasks: XNLI, KLUE-NLI, AmericasNLI, XCOPA, XStoryCloze, XWinograd, and PAWS-X.

  3. Knowl 3 — Three adaptation methods keep different parts of BLOOM trainable

    model/method

    Continued pretraining updates BLOOM with its causal language-modeling objective on monolingual text and makes the embedding layer trainable. MAD-X adds bottleneck language adapters to Transformer blocks and invertible adapters at the embedding layer; the original embeddings remain frozen. (IA)³ learns element-wise scaling vectors for Transformer activations and is paired with invertible adapters at the embedding layer, again keeping the original embeddings frozen. The invertible adapters allow the embedding representation to adjust without replacing BLOOM’s byte-level tokenizer.

  4. Knowl 4 — Prompting performance generally rises with adaptation-data quantity

    empirical result

    For BLOOM-3B adapted with 1,000, 10,000, or 100,000 samples, zero-shot prompting performance generally increased as the adaptation corpus grew. With fewer than 100,000 samples, adaptation could make performance worse than the unadapted model—for example, on Russian XNLI and Turkish XCOPA—although the low-data penalty varied by task and was limited on Russian XWinograd and XStoryCloze. The authors infer that effective adaptation in their setting required around 100 million tokens of new-language text. This data requirement also helps explain why adaptation was weak for Guarani, for which the available corpus contained only about 1 million tokens.

  5. Knowl 5 — Language family, word order, and script did not explain average XNLI results

    empirical result

    Across the tested languages, the authors found no significant difference in average XNLI prompting performance attributable to whether a language was Indo-European, used SVO rather than SOV word order, or used a script seen rather than unseen during BLOOM pretraining. This suggests that, within this experiment, those language characteristics were less predictive of prompting performance than factors such as the amount of adaptation data. The comparison covers the study’s selected languages and should not be read as a universal claim about all languages.

  6. Knowl 6 — Perplexity and sentence retrieval do not track prompting in the same way

    empirical result

    On held-out Russian text, continued-pretraining models achieved lower perplexity than MAD-X models, but this did not imply better downstream prompting: for larger BLOOM sizes, continued pretraining underperformed MAD-X on Russian XNLI and XWinograd. Cross-lingual sentence retrieval showed a different pattern. Retrieval accuracy between Russian representations from the adapted model and English representations from the original model improved with model size for MAD-X, while continued pretraining achieved its best retrieval result with BLOOM-560M and substantially lower results at larger sizes. The authors hypothesize that larger models may have more capacity to separate language representations during continued pretraining, weakening their alignment with English; this explanation is not established as a causal result.

  7. Knowl 7 — Adding adapter capacity was associated with better prompting

    empirical result

    In MAD-X adapters for BLOOM-3B, the authors varied the bottleneck reduction factor across 16, 48, and 384. A smaller reduction factor corresponds to a larger adapter bottleneck and more adapter parameters. Across the tested Russian and Turkish tasks, prompting performance generally increased with adapter capacity, giving a positive association between the number of adapter parameters and zero-shot accuracy.

  8. Knowl 8 — Monolingual adaptation of BLOOMZ can erase instruction-following performance

    empirical result

    For German XNLI, BLOOMZ-560M had a median prompt accuracy of about 38.5%, while training new adapters for BLOOMZ on free-form German OSCAR text reduced performance to about 33%, approximately the random-classifier level for the three-way task. Applying the German adapters trained on BLOOM instead preserved much of BLOOMZ’s prompting ability. Since BLOOM and BLOOMZ share an architecture, the result indicates that monolingual adaptation of the instruction-tuned model can disrupt its prompting behavior, whereas reusing adapters trained on the underlying BLOOM model avoided the same severe drop.

  9. Knowl 9 — Adding Russian to a diverse instruction-tuning mixture improved BLOOMZ

    empirical result

    The authors compared two Russian-capable 7.1-billion-parameter BLOOMZ variants: one instruction-tuned only on Russian task data, and one instruction-tuned on the full xP3 multilingual task mixture with Russian data added. Russian-only instruction tuning produced only a small improvement over the pretrained baseline on XStoryCloze. Adding Russian to the full, diverse xP3 mixture improved the best-prompt performance on Russian XNLI and XStoryCloze. The results therefore favor incorporating the new language into a broad multitask instruction-tuning mixture over tuning only on that language’s smaller set of tasks and prompts.

  10. Knowl 10 — The findings are limited by task coverage, data availability, and forgetting

    limitation

    The evaluation covered NLU prompting rather than generative tasks, so the reported adaptation results do not establish performance on tasks such as summarization. The experiments did not test vocabulary or tokenizer adaptation, and they stopped at 7.1 billion parameters rather than adapting the 176-billion-parameter BLOOM model. Only Guarani represented a truly low-resource language, and its limited adaptation data constrained the study’s evidence about such languages. The authors also observed that continued pretraining caused catastrophic forgetting on English XNLI across tested model sizes, limiting the suitability of these monolingually adapted models for tasks requiring retention of multiple languages.

Coverage note — The adapter-placement ablation and detailed per-language token-count breakdown were omitted because they are secondary to the main strategy, data-scale, and BLOOMZ findings.

References

  1. 1.Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, and Rico Sennrich. 2020. In neural machine translation, what does transfer learning transfer? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7701–7710.
  2. 2.Alan Ansell, Edoardo Ponti, Anna Korhonen, and Ivan Vulic. 2022. Composable sparse fine-tuning for cross-lingual transfer. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1778–1796, Dublin, Ireland. Association for Computational Linguistics.
  3. 3.Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, et al. 2021. Efficient large scale language modeling with mixtures of experts. arXiv preprint arXiv:2112.10684.
  4. 4.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, Online. Association for Computational Linguistics.
  5. 5.Mikel Artetxe and Holger Schwenk. 2019. Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond. Transactions of the Association for Computational Linguistics, 7:597–610.
  6. 6.Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. 2022. BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1–9, Dublin, Ireland. Association for Computational Linguistics.
  7. 7.Alexandre Berard. 2021. Continual learning in multilingual NMT via language-specific embeddings. In Proceedings of the Sixth Conference on Machine Translation, pages 542–565, Online. Association for Computational Linguistics.
  8. 8.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021. Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow. If you use this software, please cite it using these metadata, 58.
  9. 9.Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. GPT-NeoX-20B: An open-source autoregressive language model. In Proceedings of BigScience Episode #5 – Workshop on Challenges & Perspectives in Creating Large Language Models, pages 95–136, virtual+Dublin. Association for Computational Linguistics.
  10. 10.Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020. Parsing with multilingual BERT, a small corpus, and a small treebank. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1324–1334, Online. Association for Computational Linguistics.
  11. 11.Luis Chiruzzo, Santiago Góngora, Aldo Alvarez, Gustavo Giménez-Lugo, Marvin Agüero-Torales, and Yliana Rodríguez. 2022. Jojajovai: A parallel Guarani-Spanish corpus for MT benchmarking. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 2098–2107, Marseille, France. European Language Resources Association.
  12. 12.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  13. 13.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  14. 14.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  15. 15.Javier De la Rosa and Andrés Fernández. 2022. Zero-shot reading comprehension and reasoning for spanish with bertin gpt-j-6b. Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2022).
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  17. 17.Philipp Dufter and Hinrich Schütze. 2020. Identifying elements essential for BERT’s multilinguality. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4423–4437, Online. Association for Computational Linguistics.
  18. 18.Abteen Ebrahimi and Katharina Kann. 2021. How to adapt your pretrained multilingual model to 1600 languages. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4555–4567, Online. Association for Computational Linguistics.
  19. 19.Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, and Katharina Kann. 2022. AmericasNLI: Evaluating zero-shot natural language understanding of pretrained multilingual models in truly low-resource languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6279–6299, Dublin, Ireland. Association for Computational Linguistics.
  20. 20.Fahim Faisal and Antonios Anastasopoulos. 2022. Phylogeny-inspired adaptation of multilingual models to new languages. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 434–452, Online only. Association for Computational Linguistics.
  21. 21.Jonathan Frankle and Michael Carbin. 2019. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations.
  22. 22.Jinlan Fu, See-Kiong Ng, and Pengfei Liu. 2022. Polyglot prompt: Multilingual multitask promptraining. arXiv preprint arXiv:2204.14264.
  23. 23.Yoshinari Fujinuma, Jordan Boyd-Graber, and Katharina Kann. 2022. Match the script, adapt if multilingual: Analyzing the effect of multilingual pretraining on cross-lingual transferability. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1500–1512, Dublin, Ireland. Association for Computational Linguistics.
  24. 24.Philip Gage. 1994. A new algorithm for data compression. C Users Journal, 12(2):23–38.
  25. 25.Demi Guo, Alexander Rush, and Yoon Kim. 2021. Parameter-efficient transfer learning with diff pruning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4884–4896, Online. Association for Computational Linguistics.
  26. 26.Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2022. Towards a unified view of parameter-efficient transfer learning. In International Conference on Learning Representations.
  27. 27.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pages 2790–2799. PMLR.
  28. 28.Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  29. 29.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  30. 30.Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2021. Indobertweet: A pretrained language model for indonesian twitter with effective domain-specific vocabulary initialization. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10660–10668.
  31. 31.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022. Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets. Transactions of the Association for Computational Linguistics, 10:50–72.
  32. 32.Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Šaško, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa, Paulo Villegas, Tristan Thrush, Shayne Longpre, Sebastian Nagel, Leon Weber, Manuel Romero Muñoz, Jian Zhu, Daniel Van Strien, Zaid Alyafeai, Khalid Almubarak, Vu Minh Chien, Itziar Gonzalez-Dios, Aitor Soroa, Kyle Lo, Manan Dey, Pedro Ortiz Suarez, Aaron Gokaslan, Shamik Bose, David Ifeoluwa Adelani, Long Phan, Hieu Tran, Ian Yu, Suhas Pai, Jenny Chim, Violette Lepercq, Suzana Ilic, Margaret Mitchell, Sasha Luccioni, and Yacine Jernite. 2022. The bigscience ROOTS corpus: A 1.6TB composite multilingual dataset. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
  33. 33.Teven Le Scao and Alexander Rush. 2021. How many data points is a prompt worth? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2627–2636, Online. Association for Computational Linguistics.
  34. 34.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  35. 35.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas, Drew A. Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2022. Holistic evaluation of language models.
  36. 36.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2021. Few-shot learning with multilingual language models.
  37. 37.Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. arXiv preprint arXiv:2205.05638.
  38. 38.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  39. 39.Yuning Mao, Lambert Mathias, Rui Hou, Amjad Almahairi, Hao Ma, Jiawei Han, Scott Yih, and Madian Khabsa. 2022. UniPELT: A unified framework for parameter-efficient language model tuning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6253–6264, Dublin, Ireland. Association for Computational Linguistics.
  40. 40.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786.
  41. 41.Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021. When being unseen from mBERT is just the beginning: Handling new languages with multilingual language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 448–462, Online. Association for Computational Linguistics.
  42. 42.Martin Müller and Florian Laurent. 2022. Cedille: A large autoregressive french language model. arXiv preprint arXiv:2202.03371.
  43. 43.Graham Neubig and Junjie Hu. 2018. Rapid adaptation of neural machine translation to new languages. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 875–880, Brussels, Belgium. Association for Computational Linguistics.
  44. 44.NovelAI. 2022. Data efficient language transfer with gpt-j. Accessed: 2023-01-16.
  45. 45.Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. 2019. Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures. In Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-7) 2019. Cardiff, 22nd July 2019, pages 9 – 16, Mannheim. Leibniz-Institut für Deutsche Sprache.
  46. 46.Sungjoon Park, Jihyung Moon, Sungdong Kim, Won Ik Cho, Ji Yoon Han, Jangwon Park, Chisung Song, Junseong Kim, Youngsook Song, Taehwan Oh, Joohong Lee, Juhyun Oh, Sungwon Lyu, Younghoon Jeong, Inkwon Lee, Sangwoo Seo, Dongjun Lee, Hyunwoo Kim, Myeonghwa Lee, Seongbo Jang, Seungwon Do, Sunkyoung Kim, Kyungtae Lim, Jongwon Lee, Kyumin Park, Jamin Shin, Seonghyun Kim, Lucy Park, Alice Oh, Jung-Woo Ha, and Kyunghyun Cho. 2021. KLUE: Korean language understanding evaluation. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  47. 47.Marinela Parovic, Goran Glavaš, Ivan Vuli c, and Anna Korhonen. 2022. BAD-X: Bilingual adapters improve zero-shot cross-lingual transfer. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1791–1799, Seattle, United States. Association for Computational Linguistics.
  48. 48.Jonas Pfeiffer, Naman Goyal, Xi Lin, Xian Li, James Cross, Sebastian Riedel, and Mikel Artetxe. 2022. Lifting the curse of multilinguality by pre-training modular transformers. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3479–3495, Seattle, United States. Association for Computational Linguistics.
  49. 49.Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021a. AdapterFusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 487–503, Online. Association for Computational Linguistics.
  50. 50.Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder. 2020. MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7654–7673, Online. Association for Computational Linguistics.
  51. 51.Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder. 2021b. UNKs everywhere: Adapting multilingual language models to new scripts. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10186–10203, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  52. 52.Jerin Philip, Alexandre Berard, Matthias Gallé, and Laurent Besacier. 2020. Monolingual adapters for zero-shot neural machine translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4465–4470, Online. Association for Computational Linguistics.
  53. 53.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulic, and Anna Korhonen. 2020. XCOPA: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376, Online. Association for Computational Linguistics.
  54. 54.Ofir Press, Noah Smith, and Mike Lewis. 2022. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations.
  55. 55.Kunxun Qi, Hai Wan, Jianfeng Du, and Haolan Chen. 2022. Enhancing cross-lingual natural language inference by prompt-learning from cross-lingual templates. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1910–1923, Dublin, Ireland. Association for Computational Linguistics.
  56. 56.Andreas Rücklé, Gregor Geigle, Max Glockner, Tilman Beck, Jonas Pfeiffer, Nils Reimers, and Iryna Gurevych. 2021. AdapterDrop: On the efficiency of adapters in transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7930–7946, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  57. 57.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model.
  58. 58.Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725.
  59. 59.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2023. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations.
  60. 60.Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Vladislav Mikhailov, Anastasia Kozlova, and Tatiana Shavrina. 2022. mgpt: Few-shot learners go multilingual.
  61. 61.Yi-Lin Sung, Jaemin Cho, and Mohit Bansal. 2022. Lst: Ladder side-tuning for parameter and memory efficient transfer learning. arXiv preprint arXiv:2206.06522.
  62. 62.Yi-Lin Sung, Varun Nair, and Colin Raffel. 2021. Training neural networks with fixed sparse masks. In Advances in Neural Information Processing Systems.
  63. 63.Zeerak Talat, Aurélie Névéol, Stella Biderman, Miruna Clinciu, Manan Dey, Shayne Longpre, Sasha Luccioni, Maraim Masoud, Margaret Mitchell, Dragomir Radev, Shanya Sharma, Arjun Subramonian, Jaesung Tae, Samson Tan, Deepak Tunuguntla, and Oskar Van Der Wal. 2022. You reap what you sow: On the challenges of bias evaluation under multilingual settings. In Proceedings of BigScience Episode #5 – Workshop on Challenges & Perspectives in Creating Large Language Models, pages 26–41, virtual+Dublin. Association for Computational Linguistics.
  64. 64.Alexey Tikhonov and Max Ryabinin. 2021. It’s All in the Heads: Using Attention Heads as a Baseline for Cross-Lingual Transfer in Commonsense Reasoning. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3534–3546, Online. Association for Computational Linguistics.
  65. 65.Lifu Tu, Caiming Xiong, and Yingbo Zhou. 2022. Prompt-tuning can be much better than fine-tuning on cross-lingual understanding with multilingual language models. arXiv preprint arXiv:2210.12360.
  66. 66.Ahmet Üstün, Alexandre Berard, Laurent Besacier, and Matthias Gallé. 2021. Multilingual unsupervised neural machine translation with denoising adapters. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6650–6662, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  67. 67.Ben Wang and Aran Komatsuzaki. 2021. Gpt-j-6b: A 6 billion parameter autoregressive language model.
  68. 68.Zihan Wang, Karthikeyan K, Stephen Mayhew, and Dan Roth. 2020. Extending multilingual BERT to low-resource languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2649–2656, Online. Association for Computational Linguistics.
  69. 69.Genta Winata, Shijie Wu, Mayank Kulkarni, Thamar Solorio, and Daniel Preotiuc-Pietro. 2022. Cross-lingual few-shot learning on unseen languages. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 777–791, Online only. Association for Computational Linguistics.
  70. 70.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3687–3692, Hong Kong, China. Association for Computational Linguistics.
  71. 71.Zheng-Xin Yong and Vassilina Nikoulina. 2022. Adapting bigscience multilingual model to unseen languages. arXiv preprint arXiv:2204.04873.
  72. 72.Rong Zhang, Revanth Gangi Reddy, Md Arafat Sultan, Vittorio Castelli, Anthony Ferritto, Radu Florian, Efsun Sarioglu Kayi, Salim Roukos, Avi Sil, and Todd Ward. 2020. Multi-stage pre-training for low-resource domain adaptation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5461–5468, Online. Association for Computational Linguistics.
  73. 73.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  74. 74.Mengjie Zhao and Hinrich Schütze. 2021. Discrete and soft prompting for multilingual models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8547–8555, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.

Citation

MLA
Yong, Z. X., et al. “BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 11682–703, https://doi.org/10.18653/v1/2023.acl-long.653.
APA
Yong, Z. X., Schoelkopf, H., Muennighoff, N., Aji, A. F., Adelani, D. I., Almubarak, K., Bari, M. S., Sutawika, L., Kasai, J., Baruwa, A., Winata, G. I., Biderman, S., Raff, E., Radev, D., & Nikoulina, V. (2023). BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11682–11703. https://doi.org/10.18653/v1/2023.acl-long.653
Chicago
Yong, Z. X., H. Schoelkopf, N. Muennighoff, et al. 2023. “BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 11682–703. https://doi.org/10.18653/v1/2023.acl-long.653.
Harvard
Yong, Z.X. et al. (2023) “BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 11682–11703. Available at: https://doi.org/10.18653/v1/2023.acl-long.653.
Vancouver
1. Yong ZX, Schoelkopf H, Muennighoff N, et al (2023) BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 11682–11703

BibTeX

@inproceedings{yong-etal-2023-bloom,
    title = "{BLOOM}+1: Adding Language Support to {BLOOM} for Zero-Shot Prompting",
    author = "Yong, Zheng Xin  and
      Schoelkopf, Hailey  and
      Muennighoff, Niklas  and
      Aji, Alham Fikri  and
      Adelani, David Ifeoluwa  and
      Almubarak, Khalid  and
      Bari, M Saiful  and
      Sutawika, Lintang  and
      Kasai, Jungo  and
      Baruwa, Ahmed  and
      Winata, Genta  and
      Biderman, Stella  and
      Raff, Edward  and
      Radev, Dragomir  and
      Nikoulina, Vassilina",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.653/",
    doi = "10.18653/v1/2023.acl-long.653",
    pages = "11682--11703"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/