BYOL: Bring Your Own Language into LLMs
Waqas ZamirWassim HamidoucheBoulbaba Ben AmorLuana MarottiInbal Becker-ReshefJuan M. Lavista Ferres
Presents a scalable framework for integrating low- and extreme-low-resource languages into large language models, demonstrating how targeted data refinement, weight-space merging, and tailored translation pipelines significantly improve performance on underrepresented languages like Chichewa, Maori, and Inuktitut without degrading existing multilingual capabilities.
Modern artificial intelligence is heavily skewed toward a small fraction of the world’s languages, creating a pronounced digital divide. While over 7,000 languages are spoken globally, web-scale pretraining corpora are overwhelmingly dominated by English and fewer than twenty major languages. As large language models become essential infrastructure for public services, healthcare, and economic productivity, underrepresented language communities face severe performance deficits, higher inference costs, and cultural misalignment. Generic scaling of multilingual models often encounters diminishing returns and degrades existing capabilities. The article addresses this imbalance by evaluating a structured, scalable approach to integrate underrepresented languages without degrading original model quality or compromising safety.
The main objective of the article is to demonstrate Bring Your Own Language (BYOL), a resource-adaptive framework that categorizes languages by their digital footprint to determine the most effective integration pathway, ranging from direct model adaptation to translation-mediated interfaces.
The approach classifies languages into four distinct tiers based on available web text: Extreme-Low, Low, Mid, and High. For the low-resource tier, the method combines automated corpus cleaning, synthetic data generation via machine translation, continual pretraining, instruction tuning, and weight-space model merging. The article evaluates this pathway through case studies on Chichewa and Māori across a broad battery of twelve benchmarks. For the extreme-low-resource tier, the approach develops a translation-mediated pipeline that pairs neural machine translation with generalist reasoning models, tested on Inuktitut using both existing parallel text and synthetic back-translation.
The findings show substantial performance gains across all evaluated settings. First, continually pretrained and merged models (BYOL-nya and BYOL-mri) achieved an average improvement of approximately 12% over strong multilingual baselines across twelve benchmarks in Chichewa and Māori. Second, relatively compact adapted models exhibited superior efficiency: the 4-billion-parameter adapted models outperformed a standard 27-billion-parameter baseline on target language benchmarks. Third, weight-space model merging successfully preserved English performance and multilingual retention while maintaining baseline safety against toxicity and bias. Fourth, for extreme-low-resource settings, the tailored Inuktitut translation system achieved a gain of roughly 4 BLEU points over commercial baselines, and translation-mediated querying improved downstream reasoning accuracy by 11% to 15% compared to direct model inference.
These results demonstrate that language-aware adaptation is a practical alternative to massive multilingual pretraining from scratch. Organizations can deliver high-performing, localized artificial intelligence services at a fraction of the computational and financial cost typically required for large generalist models. Furthermore, model merging eliminates the need for expensive secondary safety realignment, significantly reducing implementation timelines and compliance risks for enterprise and public-sector deployments.
Decision-makers seeking to support low-resource language applications should adopt resource-tiered strategies rather than attempting direct fine-tuning indiscriminately. Teams should implement data-refinement pipelines and leverage weight merging to protect core model capabilities. Where digital data is too scarce for native training, organizations should invest in targeted translation front-ends to bridge user access to frontier reasoning models. Future initiatives should expand community-driven data collection and explore multilingual speech interfaces for traditionally oral languages.
The findings are subject to several boundary conditions, notably that native model adaptation requires a minimum threshold of usable digital text, and translation-mediated pathways remain vulnerable to cascading translation errors. Machine-translated evaluation benchmarks may also introduce subtle linguistic biases compared to fully human-annotated sets. Nevertheless, given the consistent cross-benchmark validation and robust ablations, confidence in the framework’s core methodology and operational efficiency remains high.
- Paper: The State and Fate of Linguistic Diversity and Inclusion in the NLP World, Pratik Joshi et al. (2020). It introduces the foundational taxonomy of linguistic resource disparities and digital representation across global languages that motivates BYOL's tiered integration framework.
- Paper: No Language Left Behind: Scaling Human-Centered Machine Translation, NLLB Team et al. (2022). It provides key methodology and data resources for massive multilingual scaling and low-resource machine translation that BYOL adapts for extreme-low-resource language pathways.
- Paper: Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages, Ayyoob Imani et al. (2023). It demonstrates data curation and continual pretraining recipes to horizontally expand language models to hundreds of tail languages, establishing techniques refined in BYOL's pipeline.
- Paper: Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning, Shivalika Singh et al. (2024). It establishes large-scale multilingual instruction fine-tuning pipelines and evaluation suites for low-resource languages that BYOL builds upon.
- Paper: Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model, Ahmet Üstün et al. (2024). It provides practical recipes for mixing human-curated, translated, and synthetic instruction datasets to adapt foundation models across diverse language tiers.
- Paper: The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants, Lucas Bandarkar et al. (2024). It details parallel multilingual reading comprehension benchmarking across diverse resource levels, informing BYOL's multilingual evaluation approach.
- Paper: Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages, Zihao Li et al. (2025). It develops quantitative metrics to evaluate LLM internal representation gaps across high- and low-resource languages, providing background for BYOL's language-aware resource categorization.
- Paper: MEGA: Multilingual Evaluation of Generative AI, Kabir Ahuja et al. (2023). It benchmarks the performance gaps between native prompting and translation-mediated workflows in generative LLMs across diverse languages.
- Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). It scales translation-mediated inclusion and multilingual LLM specialization to over 1,600 languages, extending the translation pathways explored in BYOL for extreme-low-resource settings.
