BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting
Zheng Xin YongHailey SchoelkopfNiklas MuennighoffAlham Fikri AjiDavid Ifeoluwa AdelaniKhalid AlmubarakM. Saiful BariLintang SutawikaJungo KasaiAhmed Baruwa
Demonstrates that parameter-efficient adapter finetuning outperforms continued pretraining when adapting large multilingual models like BLOOM to unseen languages in resource-constrained, zero-shot prompting settings.
Large multilingual language models offer powerful capabilities, but training them from scratch requires immense computational resources. As a result, even leading models cover only a fraction of the world’s languages, excluding major populations and low-resource language communities. Because retraining entire models to add new languages is financially prohibitive, organizations need efficient techniques to expand language coverage without incurring unsustainable computing costs.
The article systematically evaluates different lightweight strategies to adapt the open-access BLOOM model family to eight previously unsupported languages. Its primary goal is to determine the most effective and computationally efficient methods for enabling zero-shot task performance in new languages under resource-constrained training budgets.
The authors conducted experimental benchmarks adapting BLOOM models ranging in scale from 560 million to 7.1 billion parameters across eight languages representing diverse scripts and families: Bulgarian, German, Greek, Guarani, Korean, Russian, Thai, and Turkish. The evaluation compared continued causal language pretraining against parameter-efficient modular approaches, primarily bottleneck adapters (MAD-X) and activation-scaling adapters ((IA)3). Training utilized a low-resource budget capped at 100,000 text samples per language (about 100 million tokens), without altering the base model’s subword vocabulary. Performance was measured across standard natural language understanding tasks, including inference, reasoning, and paraphrase detection, using zero-shot prompting.
The investigation produced several key findings regarding model scaling, architectural efficiency, and data requirements. First, adapter-based adaptation consistently outperforms continued pretraining for models with 3 billion or more parameters, while requiring substantially lower GPU memory and training runtime. Continued pretraining proved superior only for the smallest 560-million-parameter model. Second, the effectiveness of zero-shot prompting depends heavily on the volume of adaptation data rather than linguistic characteristics; unseen writing systems and diverse word orders adapted equally well, provided sufficient text was available. Third, effective adaptation requires approximately 100 million tokens; performance collapsed when training on 10,000 or fewer samples, explaining the poor results observed on the low-resource language Guarani (30,000 samples). Fourth, adapter capacity matters, as allocating more parameters to adapter modules steadily boosted task accuracy. Finally, for instruction-tuned model variants (BLOOMZ), directly incorporating target-language task data into the broader multitask training mixture proved far more effective than adapting the model solely on raw, unstructured monolingual text.
These results demonstrate that expanding large language models into new languages is viable without the extreme expense of full-model retraining. Relying on modular adapters significantly lowers compute infrastructure costs, shortens deployment timelines, and avoids catastrophic forgetting of original languages. This provides a practical path for enterprises and public institutions to democratize language technology across underrepresented languages.
Decision-makers seeking to extend large language models should deploy modular adapter frameworks rather than full continued pretraining when working with models of 3 billion parameters or larger. When deploying instruction-following assistants, teams should integrate target-language examples directly into diverse multitask fine-tuning mixtures. Projects should ensure a baseline training corpus of roughly 100 million tokens per language before attempting adaptation. If only small data volumes are available, organizations should run pilot benchmarks first, as low-data regimes can degrade baseline performance.
Confidence in these findings is high for classification and reasoning tasks across small- to medium-sized model tiers. However, caution is warranted when extrapolating these conclusions to text-generation tasks or extremely low-resource languages lacking sufficient text corpora. Additionally, computational constraints prevented testing on the maximum 176-billion-parameter BLOOM architecture, which remains an area for future validation.
- Paper: BLOOM: A 176B-Parameter Open-Access Multilingual Language Model, BigScience Workshop (2022). Read the BLOOM model paper first to understand the model’s architecture, training data, and 46-language coverage that this work seeks to extend.
- Paper: mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Linting Xue et al. (2020). The mT5 paper establishes multilingual pretraining and zero-shot evaluation practices that frame this study’s work on adapting large pretrained models across languages.
- Paper: Multitask Prompted Training Enables Zero-Shot Task Generalization, Victor Sanh et al. (2021). T0 shows how multitask prompted training produces zero-shot task generalization, providing essential context for this paper’s BLOOMZ language-adaptation experiments.
- Paper: How Multilingual is Multilingual BERT?, Telmo Pires et al. (2019). This study’s evidence for cross-lingual zero-shot transfer clarifies the multilingual capabilities and limitations that motivate adapting BLOOM to unseen languages.
- Paper: Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation, Xinyi Wang et al. (2022). Its lexicon-based adaptation strategies offer a concrete precedent for extending pretrained models to languages beyond their original coverage.
- Paper: Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model, Ahmet Üstün et al. (2024). Aya carries multilingual instruction finetuning further, scaling open-access instruction following to 101 languages and enabling comparison with BLOOMZ’s language-expansion approach.
- Paper: Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning, Shivalika Singh et al. (2024). The Aya Dataset develops multilingual instruction-tuning resources at much larger language coverage, extending the source’s finding that target-language examples improve instruction following.
- Paper: FinGPT: Large Generative Models for a Small Language, Risto Luukkonen et al. (2023). FinGPT applies continued pretraining to BLOOM for Finnish, providing a later language-specific case study to compare with the source’s adaptation strategies.
- Paper: The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants, Lucas Bandarkar et al. (2024). BELEBELE tests multilingual models across 122 language variants, extending the source’s focus on evaluating language coverage and performance beyond major languages.
