BYOL: Bring Your Own Language into LLMs

Waqas ZamirWassim HamidoucheBoulbaba Ben AmorLuana MarottiInbal Becker-ReshefJuan M. Lavista Ferres

article2026arXiv0 citations

Presents a scalable framework for integrating low- and extreme-low-resource languages into large language models, demonstrating how targeted data refinement, weight-space merging, and tailored translation pipelines significantly improve performance on underrepresented languages like Chichewa, Maori, and Inuktitut without degrading existing multilingual capabilities.

Listen

Modern artificial intelligence is heavily skewed toward a small fraction of the world’s languages, creating a pronounced digital divide. While over 7,000 languages are spoken globally, web-scale pretraining corpora are overwhelmingly dominated by English and fewer than twenty major languages. As large language models become essential infrastructure for public services, healthcare, and economic productivity, underrepresented language communities face severe performance deficits, higher inference costs, and cultural misalignment. Generic scaling of multilingual models often encounters diminishing returns and degrades existing capabilities. The article addresses this imbalance by evaluating a structured, scalable approach to integrate underrepresented languages without degrading original model quality or compromising safety.

The main objective of the article is to demonstrate Bring Your Own Language (BYOL), a resource-adaptive framework that categorizes languages by their digital footprint to determine the most effective integration pathway, ranging from direct model adaptation to translation-mediated interfaces.

The approach classifies languages into four distinct tiers based on available web text: Extreme-Low, Low, Mid, and High. For the low-resource tier, the method combines automated corpus cleaning, synthetic data generation via machine translation, continual pretraining, instruction tuning, and weight-space model merging. The article evaluates this pathway through case studies on Chichewa and Māori across a broad battery of twelve benchmarks. For the extreme-low-resource tier, the approach develops a translation-mediated pipeline that pairs neural machine translation with generalist reasoning models, tested on Inuktitut using both existing parallel text and synthetic back-translation.

The findings show substantial performance gains across all evaluated settings. First, continually pretrained and merged models (BYOL-nya and BYOL-mri) achieved an average improvement of approximately 12% over strong multilingual baselines across twelve benchmarks in Chichewa and Māori. Second, relatively compact adapted models exhibited superior efficiency: the 4-billion-parameter adapted models outperformed a standard 27-billion-parameter baseline on target language benchmarks. Third, weight-space model merging successfully preserved English performance and multilingual retention while maintaining baseline safety against toxicity and bias. Fourth, for extreme-low-resource settings, the tailored Inuktitut translation system achieved a gain of roughly 4 BLEU points over commercial baselines, and translation-mediated querying improved downstream reasoning accuracy by 11% to 15% compared to direct model inference.

These results demonstrate that language-aware adaptation is a practical alternative to massive multilingual pretraining from scratch. Organizations can deliver high-performing, localized artificial intelligence services at a fraction of the computational and financial cost typically required for large generalist models. Furthermore, model merging eliminates the need for expensive secondary safety realignment, significantly reducing implementation timelines and compliance risks for enterprise and public-sector deployments.

Decision-makers seeking to support low-resource language applications should adopt resource-tiered strategies rather than attempting direct fine-tuning indiscriminately. Teams should implement data-refinement pipelines and leverage weight merging to protect core model capabilities. Where digital data is too scarce for native training, organizations should invest in targeted translation front-ends to bridge user access to frontier reasoning models. Future initiatives should expand community-driven data collection and explore multilingual speech interfaces for traditionally oral languages.

The findings are subject to several boundary conditions, notably that native model adaptation requires a minimum threshold of usable digital text, and translation-mediated pathways remain vulnerable to cascading translation errors. Machine-translated evaluation benchmarks may also introduce subtle linguistic biases compared to fully human-annotated sets. Nevertheless, given the consistent cross-benchmark validation and robust ablations, confidence in the framework’s core methodology and operational efficiency remains high.

  • Paper: Omnilingual MT: Machine Translation for 1,600 Languages, Omnilingual MT Team et al. (2026). It scales translation-mediated inclusion and multilingual LLM specialization to over 1,600 languages, extending the translation pathways explored in BYOL for extreme-low-resource settings.
Cover for BYOL: Bring Your Own Language into LLMs

Abstract

Large Language Models (LLMs) exhibit strong multilingual capabilities, yet remain fundamentally constrained by the severe imbalance in global language resources. While over 7,000 languages are spoken worldwide, only a small subset (fewer than 100) has sufficient digital presence to meaningfully influence modern LLM training. This disparity leads to systematic underperformance, cultural misalignment, and limited accessibility for speakers of low-resource and extreme-low-resource languages. To address this gap, we introduce Bring Your Own Language (BYOL), a unified framework for scalable, language-aware LLM development tailored to each language's digital footprint. BYOL begins with a language resource classification that maps languages into four tiers (Extreme-Low, Low, Mid, High) using curated web-scale corpora, and uses this classification to select the appropriate integration pathway. For low-resource languages, we propose a full-stack data refinement and expansion pipeline that combines corpus cleaning, synthetic text generation, continual pretraining, and supervised finetuning. Applied to Chichewa and Maori, this pipeline yields language-specific LLMs that achieve approximately 12 percent average improvement over strong multilingual baselines across 12 benchmarks, while preserving English and multilingual capabilities via weight-space model merging. For extreme-low-resource languages, we introduce a translation-mediated inclusion pathway, and show on Inuktitut that a tailored machine translation system improves over a commercial baseline by 4 BLEU, enabling high-accuracy LLM access when direct language modeling is infeasible. Finally, we release human-translated versions of the Global MMLU-Lite benchmark in Chichewa, Maori, and Inuktitut, and make our codebase and models publicly available at this https URL .

Table of Contents

  • 1 Introduction
  • 2 Bring Your Own Language (BYOL) Framework
  • 2.1 Initial Assessment
  • 2.1.1 Language Resource Classification
  • 2.1.2 Existing Tools Evaluation: LLMs and MT Systems
  • 2.2 Low-Resource Pathway: Native Language Support in LLMs
  • 2.3 Extreme-Low-Resource Pathway: LLM Access via Machine Translation
  • 3 Experiments and Analysis
  • 3.1 Direct LLM Adaptation for Chichewa and Māori
  • 3.1.1 Experimental Details
  • 3.1.2 Performance Evaluation
  • 3.1.3 Ablation Studies
  • 3.2 Translation-Mediated LLM Access for Inuktitut
  • 3.2.1 Experimental Details
  • 3.2.2 Performance Evaluation
  • 4 Challenges and Future Work
  • 4.1 LLM Safety in Low-Resource Languages
  • 4.2 Extension to Multilingual LLMs
  • 4.3 Extension to Speech
  • 4.4 Data Scarcity
  • 5 Conclusion
  • 6 Acknowledgments
  • References
  • A Related Work
  • B Training Datasets
  • B.1 Continual Pretraining Datasets
  • B.2 Instruction Finetuning Datasets
  • B.3 Machine Translation Datasets
  • C Evaluation Datasets
  • C.1 Base Model Benchmarking Datasets
  • C.2 Instruction-Tuned Model Benchmarking Datasets
  • C.3 LLM as Judge Evaluation
  • D Performance of BYOL Models on English
  • E Ablation Experiment Results
  • E.1 RTTBench-Mono Validation
  • E.2 Evaluation of MTs on RTT Benchmarks
  • E.3 Impact of CPT Data Mixture
  • E.4 Ablation of LoRA vs. Full-Parameter
  • E.5 Impact of Model Merging Weight
  • F Prompts
  • F.1 Prompt for Text Refinement
  • F.2 Prompt for Sentence Alignment in Extreme-Low-Resource Languages
  • F.3 Prompt for LLM-Based Enhancement of English Text for Back-Translation
  • F.4 Prompt for LLM-Based Post-Editing of MT Outputs
  • F.5 Prompt to Generate RTTBench-Mono
  • F.6 Prompt Template for LLM-as-a-Judge Evaluation

Knowls

  1. Knowl 1 — Four-Tier Language Resource Classification and Routing Scheme

    model/method

    The Bring Your Own Language (BYOL) framework stratifies human languages into four discrete corpus-size tiers based on the total deduplicated word count available in web-scale multilingual corpora (specifically FineWeb2):

    1. Extreme-Low-Resource (le5×106\\le 5 \times 10^6 words): Languages with negligible digital presence. Direct pretraining or finetuning produces unintelligible outputs. The framework routes these languages through a Translate-Test pathway, utilizing a dedicated machine translation (MT) system to interface with high-resource LLMs.
    2. Low-Resource (5×106−2×1095 \times 10^6 - 2 \times 10^9 words): Languages with limited, noisy, but usable corpora. The framework routes these through a Native Language Support pathway involving data refinement, synthetic text generation, continual pretraining (CPT), supervised finetuning (SFT), and weight-space model merging.
    3. Mid-Resource (2×109−10112 \times 10^9 - 10^{11} words): Languages with substantial textual resources and moderate-to-strong existing coverage, which can be adapted via targeted instruction finetuning.
    4. High-Resource (>1011> 10^{11} words): Languages with abundant web-scale corpora that already possess broad LLM support across tasks.
  2. Knowl 2 — Weight-Space Model Merging for Low-Resource Language Specialization

    equation

    To combine language-specific proficiency with general multilingual instruction-following and safety behaviors, BYOL merges model weights directly in parameter space without requiring joint retraining. Let GPTG_{\text{PT}} and GITG_{\text{IT}} represent the pretrained base and instruction-tuned variants of a multilingual generalist model, respectively, and let EℓE_\ell represent the language-specialist expert model produced by applying continual pretraining (CPT) and supervised finetuning (SFT) starting from GPTG_{\text{PT}}.

    The merged model weights M(α,β)M(\alpha, \beta) are computed as:

    M(α,β)=GPT+α(GIT−GPT)+β(Eℓ−GPT)M(\alpha, \beta) = G_{\text{PT}} + \alpha(G_{\text{IT}} - G_{\text{PT}}) + \beta(E_\ell - G_{\text{PT}})

    where α≥0\alpha \ge 0 and β≥0\beta \ge 0 are scaling coefficients. The task vector (GIT−GPT)(G_{\text{IT}} - G_{\text{PT}}) transfers general instruction-following and safety behavior, while (Eℓ−GPT)(E_\ell - G_{\text{PT}}) injects low-resource language expertise.

    In a 1D reparameterization where α=1−λ\alpha = 1 - \lambda and β=λ\beta = \lambda for λ∈[0,1]\lambda \in [0, 1], target-language proficiency improves with increasing λ\lambda while general English/multilingual proficiency gradually decreases; setting λ≈0.6\lambda \approx 0.6 achieves the optimal bilingual performance trade-off.

  3. Knowl 3 — Domain-Conditioned Round-Trip Translation (RTT) Proxy Evaluation

    model/method

    To evaluate LLMs and machine translation engines on low-resource languages lacking parallel gold benchmarks, BYOL uses a domain-conditioned Round-Trip Translation (RTT) framework. A set of sentences from a high-resource pivot language (e.g., English) is translated into the target language ℓ\ell and back to English, comparing the reconstructed output with the original source.

    The metric is defined as:

    RTTScore=1∣D∣∑d∈D1Nd∑i=1NdM(si(d),Tℓ→eng(Teng→ℓ(si(d))))\text{RTTScore} = \frac{1}{|D|} \sum_{d \in D} \frac{1}{N_d} \sum_{i=1}^{N_d} M\left(s_i^{(d)}, T_{\ell \to \text{eng}}\left(T_{\text{eng} \to \ell}\left(s_i^{(d)}\right)\right)\right)

    where DD denotes the set of domain categories, NdN_d is the number of sentences sampled in domain dd, si(d)s_i^{(d)} is the ii-th sentence from domain dd, Teng→ℓT_{\text{eng} \to \ell} and Tℓ→engT_{\ell \to \text{eng}} denote the forward and back-translation models, and MM is a text fidelity metric (such as SacreBLEU, chrF++, or text embedding cosine similarity).

    This is operationalized on RTTBench-Mono, a domain-balanced English dataset containing 1,250 sentences evenly distributed across 25 distinct topical domains (50 sentences per domain) with controlled length quotas and syntax variation.

  4. Knowl 4 — Data Refinement and Bilingual Continual Pretraining Recipe for LRLs

    model/method

    For low-resource languages (LRLs), BYOL builds a balanced continual pretraining (CPT) corpus using a 1:1:1 token ratio across three components:

    1. Refined Native LRL Text: Real text filtered from FineWeb2, rewritten and enriched via guided rephrasing using an LLM (e.g., GPT-5) to repair grammar, improve coherence, remove toxic content, and expand length to 100–140% of the input without shortening.
    2. Synthetic LRL Text: High-quality English educational text (FineWeb-Edu) refined by an LLM, translated into the target language ℓ\ell using the top-performing MT engine identified via RTT evaluation.
    3. Refined English Text: A subset of refined FineWeb-Edu English samples included to prevent catastrophic forgetting of base English and cross-lingual reasoning skills.

    Models are trained on this mixture using AdamW (β1=0.9,β2=0.999\beta_1 = 0.9, \beta_2 = 0.999) with cosine learning rate decay (2×10−5→2×10−62 \times 10^{-5} \to 2 \times 10^{-6}), 3% linear warm-up, a context length of 4096 tokens, for 4 epochs using full-parameter training.

  5. Knowl 5 — Performance of Direct LLM Adaptation on Chichewa and Māori Benchmarks

    data/table

    BYOL adaptation was evaluated on two typologically diverse low-resource languages: Chichewa (Bantu family) and Māori (Austronesian family). Models were tested across 12 benchmarks: Global MMLU-Lite, ARC-Hard chat, MGSM, XCOPA, XStoryCloze, PIQA, HellaSwag, XNLI 2.0, XWinograd, Belebele, FLORES-200 (translation), and TruthfulQA.

    Benchmark Metric Apertus Qwen-3 Gemma-3 BYOL (Ours) Gemma-3 BYOL (Ours)
    (8B-Inst) (8B) (4B-IT) (4B-M) (27B-IT) (12B-M)
    Chichewa (nya)
    Global MMLU-Lite (Acc) 34.83 36.45 45.36 53.62 62.64 66.15
    ARC-Hard chat (Acc) 32.51 17.66 33.28 50.43 66.13 64.76
    MGSM (Acc) 2.40 2.40 11.20 30.00 38.80 40.80
    Belebele (Acc) 34.56 22.11 29.00 55.00 52.67 62.44
    FLORES nya→\toeng (BLEU) 11.53 6.11 11.97 24.96 22.84 27.13
    Average Score (12 tasks) 35.69 31.66 36.91 50.89 49.90 55.47
    Māori (mri)
    Global MMLU-Lite (Acc) 31.05 33.27 35.10 47.64 54.64 52.48
    ARC-Hard chat (Acc) 27.30 34.04 37.50 51.11 49.66 59.22
    MGSM (Acc) 4.80 7.20 10.80 27.60 41.60 38.80
    Belebele (Acc) 33.22 22.89 25.56 50.67 50.33 62.78
    FLORES mri→\toeng (BLEU) 14.22 11.83 11.64 26.02 21.70 28.14
    Average Score (12 tasks) 36.68 36.08 36.97 48.38 45.52 52.35

    The 4B merged models (BYOL-nya 4B-M and BYOL-mri 4B-M) outperform the ~7×\times larger Gemma-3 27B-IT baseline across both languages by approximately 1% to 3% on average, and exceed same-sized 4B–8B baselines by roughly 12% to 15%.

  6. Knowl 6 — Safety and Multilingual Capability Retention via Weight Merging

    empirical result

    Direct instruction finetuning on low-resource datasets degrades base model safety guardrails and cross-lingual performance on non-target languages. Applying weight-space model merging (M(α,β)M(\alpha, \beta)) restores both capabilities while preserving low-resource gains:

    1. Safety Restoration: On standard safety benchmarks (BBQ for bias, ToxiGen and RealToxicityPrompts for toxic generation), unmerged models (e.g., BYOL-nya (IT)) exhibit severe safety degradation, with RealToxicityPrompts toxicity scores rising from 0.35 (Gemma-3 4B baseline) to 4.79–5.87. Merged models (BYOL-nya (M)) restore the toxicity metric to 1.44–1.53 without additional alignment data.
    2. Multilingual Retention: On Global MMLU-Lite evaluated over 17 non-target languages (excluding English and the target language), the unmerged 4B IT model suffers an average accuracy drop from 49.0% to 47.0%, whereas the merged 4B-M model recovers and improves the average non-target multilingual score to 49.3%.
  7. Knowl 7 — Translation-Mediated Pathway for Extreme-Low-Resource Languages

    model/method

    For extreme-low-resource languages lacking sufficient text for direct LLM training, BYOL implements a modular translation-mediated pipeline consisting of three stages:

    1. LLM-Based Sentence Alignment: Rather than relying on length heuristics or embedding similarity models (which fail on extreme LRLs), a multilingual LLM (such as GPT-5-chat) detects monotonic cross-lingual correspondences directly from unaligned bilingual documents.
    2. Synthetic Data Expansion via Back-Translation: A preliminary bidirectional NMT system is trained on available bitext and used to back-translate clean monolingual English and target-language corpora. Real and back-translated parallel pairs are mixed in a 1:1 ratio.
    3. LLM-Based Post-Editing: Raw NMT outputs are refined using an LLM instructed to perform minimal corrections on lexical choices, grammar, and fluency without altering sentence structure or hallucinating facts.
    4. Translate-Test Inference: User input in the extreme LRL ℓ\ell is translated into English (Tℓ→engT_{\ell \to \text{eng}}), processed by an English LLM, and translated back (Teng→ℓT_{\text{eng} \to \ell}) to provide the final native response.
  8. Knowl 8 — Inuktitut Translation Performance and LLM Accuracy Recovery

    empirical result

    Applying the translation-mediated inclusion strategy to Inuktitut (ISO 639-3: iku) demonstrated superior machine translation accuracy and substantial downstream LLM performance recovery:

    • NMT Quality: The BYOL Inuktitut Transformer models (9-layer encoder, 9-layer decoder, 4096-token shared BPE vocabulary) achieved 32.80 BLEU / 51.22 chrF++ on Inuktitut→\toEnglish (+3.64 BLEU over Azure Translator) and 16.25 BLEU / 45.70 chrF++ on English→\toInuktitut (+4.31 BLEU over Azure Translator) averaged across Nunavut Hansard 3.0, news articles, and children's books.
    • Downstream Reasoning Recovery: On the human-translated Global MMLU-Lite benchmark, directly prompting frontier LLMs in Inuktitut caused accuracy to drop below 35% for GPT-4o and GPT-4.1, and to 46.4% for GPT-5-Reasoning. Utilizing the translation-mediated interface (Inuktitut →\to English →\to LLM) recovered accuracy to 49.3% (+14.6% gain) for GPT-4o, 48.3% (+14.4% gain) for GPT-4.1, and 57.9% (+11.5% gain) for GPT-5-Reasoning.
  9. Knowl 9 — Full-Parameter Tuning Superiority over LoRA in LRL Continual Pretraining

    empirical result

    In continual pretraining experiments on the low-resource language Chichewa using a 4B parameter model (BYOL-nya 4B-CPT), low-rank adaptation (LoRA) scaled with rank r∈{64,128,256,512}r \in \{64, 128, 256, 512\} but consistently underperformed full-parameter tuning:

    • Gemma-3 (4B-PT) baseline average score on Chichewa benchmarks: 39.95
    • LoRA r=64r=64: 43.59 (Chichewa avg)
    • LoRA r=128r=128: 47.03 (Chichewa avg)
    • LoRA r=256r=256: 47.90 (Chichewa avg)
    • LoRA r=512r=512: 49.60 (Chichewa avg)
    • Full-parameter tuning: 51.82 (Chichewa avg)

    Full-parameter updating achieved a 2.22 point boost over the largest LoRA rank (r=512r=512) while maintaining identical English performance (65.29 vs 64.96), leading the framework to adopt full-parameter training for all CPT stages.

  10. Knowl 10 — Ablation of Continual Pretraining Data Mixture Configurations

    data/table

    Ablation on continual pretraining data compositions for BYOL-nya (4B-CPT) demonstrates the necessity of combining refined real text, synthetic translated text, and anchor English text.

    Corpus Component C1 C2 C3 C4
    Raw Native Text (FineWeb2) ✓
    Refined Native Text ✓ ✓ ✓
    Refined English Text (FineWeb-Edu) ✓ ✓
    Synthetic LRL (Refined English →\to MT) ✓
    Chichewa Average Score 48.77 49.44 49.60 51.82
    English Average Score 64.58 65.24 65.37 65.29

    The full combination (C4) outperforms raw native text alone (C1) by +3.05 points on target Chichewa tasks while avoiding catastrophic forgetting on English benchmarks.

Coverage note — Omitted the full text transcripts of the prompt templates (Appendices F.1–F.6) and per-domain bar charts, as their operational mechanisms and quantitative summaries are fully incorporated into the respective knowls.

References

  1. 1.L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al., “The pile: An 800gb dataset of diverse text for language modeling,” arXiv:2101.00027, 2020.
  2. 2.M. Weber, D. Fu, Q. Anthony, Y. Oren, S. Adams, A. Alexandrov, X. Lyu, H. Nguyen, X. Yao, V. Adams, et al., “‘RedPajama: an open dataset for training large language models,” Advances in neural information processing systems, 2024.
  3. 3.J. Abadji, P. Ortiz Suarez, L. Romary, and B. Sagot, “‘Towards a cleaner document-oriented multilingual crawled corpus,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference, 2022.
  4. 4.D. Soboleva, F. Al-Khateeb, R. Myers, J. R. Steeves, J. Hestness, and N. Dey, “‘SlimPajama: a 627b token cleaned and deduplicated version of redpajama,” Cerebras Blog, 2023.
  5. 5.G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, H. Alobeidli, A. Cappelli, B. Pannier, E. Almazrouei, and J. Launay, “‘The RefinedWeb dataset for Falcon LLM: Outperforming curated corpora with web data only,” in Advances in Neural Information Processing Systems, 2024.
  6. 6.L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, and Others, “‘Dolma: an open corpus of three trillion tokens for language model pretraining research,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024.
  7. 7.G. Penedo, H. Kydlíćek, A. Lozhkov, M. Mitchell, C. A. Raffel, L. Von Werra, T. Wolf, et al., “The fineweb datasets: Decanting the web for the finest text data at scale,” Advances in Neural Information Processing Systems, 2024.
  8. 8.P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury, “‘The state and fate of linguistic diversity and inclusion in the NLP world,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020.
  9. 9.G. Penedo, H. Kydlíćek, V. Sabolćec, B. Messmer, N. Foroutan, A. H. Kargaran, C. Raffel, M. Jaggi, L. Von Werra, and T. Wolf, “‘FineWeb2: One pipeline to scale them all–adapting pretraining data processing to every language,” arXiv:2506.20920, 2025.
  10. 10.“The ai language gap: Considerations on the multilingual capabilities of ai language models,” tech. rep., Cohere Labs, 2024. Policy Primer.
  11. 11.A. Peppin, J. Kreutzer, A. S. Sebag, K. Marchisio, B. Ermis, J. Dang, S. Cahyawijaya, S. Singh, S. Goldfarb-Tarrant, V. Aryabumi, et al., “‘The multilingual divide and its impact on global ai safety,” arXiv:2505.21344, 2025.
  12. 12.A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al., “‘Qwen3 technical report,” arXiv:2505.09388, 2025.
  13. 13.G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière, et al., “‘Gemma 3 technical report,” arXiv:2503.19786, 2025.
  14. 14.A. Hernández-Cano, A. Hägele, A. H. Huang, A. Romanou, A.-J. Solergibert, B. Pasztor, B. Messmer, D. Garbaya, E. F. Durech, I. Hakimi, et al., “‘Apertus: Democratizing open and compliant llms for global language environments,” arXiv:2509.14233, 2025.
  15. 15.A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al., “The Llama 3 herd of models,” arXiv:2407.21783, 2024.
  16. 16.S. Agarwal, L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y. Bai, B. Baker, H. Bao, et al., “‘gpt-oss-120b & gpt-oss-20b model card,” arXiv preprint arXiv:2508.10925, 2025.
  17. 17.A. Misra, J. Wang, S. McCullers, K. White, and J. L. Ferres, “‘Measuring AI diffusion: A population-normalized metric for tracking global ai usage,” arXiv:2511.02781, 2025.
  18. 18.A. Misra, S. W. Zamir, W. Hamidouche, I. Becker-Reshef, and J. Lavista Ferres, “‘AI diffusion in low resource language countries,” arXiv:2511.02752, 2025.
  19. 19.R. Appel, P. McCrory, A. Tamkin, M. McCain, T. Neylon, and M. Stern, “‘The anthropic economic index report: Uneven geographic and enterprise ai adoption,” tech. rep., Anthropic, 2025. Accessed: 2025-10-06.
  20. 20.T. Patwardhan, R. Dias, E. Proehl, G. Kim, M. Wang, O. Watkins, S. P. Fishman, M. Aljubbeh, P. Thacker, L. Fauconnet, N. S. Kim, P. Chao, S. Miserendino, G. Chabot, D. Li, M. Sharman, A. Barr, A. Glaese, and J. Tworek, “‘GDPVal: Evaluating ai model performance on real-world economically valuable tasks,” tech. rep., OpenAI, 2025. Accessed: 2025-10-06.
  21. 21.S. Singh, A. Romanou, C. Fourrier, D. I. Adelani, J. G. Ngui, D. Vila-Suero, P. Limkonchotiwat, K. Marchisio, W. Q. Leong, Y. Susanto, et al., “‘Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation,” arXiv:2412.03304, 2024.
  22. 22.C. Holtermann, P. Röttger, T. Dill, and A. Lauscher, “‘Evaluating the elementary multilingual capabilities of large language models with MultiQ,” in Findings of the Association for Computational Linguistics: ACL 2024, 2024.
  23. 23.P. Maini, V. Dorna, P. Doshi, A. Carranza, F. Pan, J. Urbanek, P. Burstein, A. Fang, A. Deng, A. Abbas, et al., “‘BeyondWeb: lessons from scaling synthetic data for trillion-scale pretraining,” arXiv:2508.10975, 2025.
  24. 24.W. Xuan, R. Yang, H. Qi, Q. Zeng, Y. Xiao, and Others, “‘MMLU-ProX: a multilingual benchmark for advanced large language model evaluation,” arXiv:2503.10497, 2025.
  25. 25.S. Ahuja, D. Aggarwal, V. Gumma, I. Watts, A. Sathe, M. Ochieng, R. Hada, P. Jain, M. Ahmed, K. Bali, et al., “‘MEGAVERSE: Benchmarking large language models across languages, modalities, models and tasks,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024.
  26. 26.T. A. Chang, C. Arnett, Z. Tu, and B. Bergen, “‘When is multilinguality a curse? language modeling for 250 high- and low-resource languages,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024.
  27. 27.N. Arivazhagan, A. Bapna, O. Firat, D. Lepikhin, M. Johnson, M. Krikun, M. X. Chen, Y. Cao, G. F. Foster, C. Cherry, W. Macherey, Z. Chen, and Y. Wu, “‘Massively multilingual neural machine translation in the wild: Findings and challenges,” arXiv:1907.05019, 2019.
  28. 28.A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “‘Unsupervised cross-lingual representation learning at scale,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020.
  29. 29.J. Pfeiffer, N. Goyal, X. V. Lin, X. Li, J. Cross, S. Riedel, and M. Artetxe, “‘Lifting the curse of multilinguality by pre-training modular transformers,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 2022.
  30. 30.A. Üstün, V. Aryabumi, Z. Yong, W.-Y. Ko, D. D’souza, G. Onilude, N. Bhandari, S. Singh, H.-L. Ooi, A. Kayid, F. Vargus, P. Blunsom, S. Longpre, N. Muennighoff, M. Fadaee, J. Kreutzer, and S. Hooker, “‘Aya Model: an instruction finetuned open-access multilingual language model,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024.
  31. 31.D. I. Adelani, J. Ojo, I. A. Azime, J. Y. Zhuang, J. O. Alabi, X. He, M. Ochieng, S. Hooker, A. Bukula, E.-S. A. Lee, C. I. Chukwuneke, H. Buzaaba, B. K. Sibanda, G. K. Kalipe, J. Mukiibi, S. Kabongo Kabenamualu, F. Yuehgoh, M. Setaka, L. Ndolela, N. Odu, R. Mabuya, S. Osei, S. H. Muhammad, S. Samb, T. K. Guge, T. V. Sherman, and P. Stenetorp, “‘IrokoBench: A new benchmark for African languages in the age of large language models,” in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025.
  32. 32.O. Ahia, S. Kumar, H. Gonen, J. Kasai, D. Mortensen, N. Smith, and Y. Tsvetkov, “‘Do all languages cost the same? tokenization in the era of commercial language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2023.
  33. 33.C. Arnett and B. Bergen, “Why do language models perform worse for morphologically complex languages?,” in Proceedings of the 31st International Conference on Computational Linguistics, 2025.
  34. 34.C. Arnett, “There is no such thing as a tokenizer-free lunch.” https://huggingface.co/blog/catherinearnett/in-defense-of-tokenizers, 2025. Hugging Face Community Article.
  35. 35.D. Abagyan, A. R. Salamanca, A. F. Cruz-Salinas, K. Cao, H. Lin, A. Locatelli, M. Fadaee, A. Üstün, and S. Hooker, “‘One tokenizer to rule them all: Emergent language plasticity via multilingual tokenizers,” arXiv:2506.10766, 2025.
  36. 36.Y. Deng, W. Zhang, S. J. Pan, and L. Bing, “‘Multilingual jailbreak challenges in large language models,” in Proceedings of the International Conference on Learning Representations (ICLR), 2024.
  37. 37.Z.-X. Yong, C. Menghini, and S. H. Bach, “‘Low-resource languages jailbreak gpt-4,” arXiv:2310.02446, 2023.
  38. 38.OECD, “‘Ai language models: Technological, socio-economic and policy considerations,” Tech. Rep. 352, Organisation for Economic Co-operation and Development, 2023.
  39. 39.W. Nekoto, V. Marivate, T. Matsila, T. Fasubaa, T. Fagbohungbe, S. O. Akinola, S. Muhammad, S. Kabongo Kabenamualu, S. Osei, F. Sackey, R. A. Niyongabo, R. Macharm, P. Ogayo, O. Ahia, M. M. Berhe, M. Adeyemi, M. Mokgesi-Selinga, L. Okegbemi, L. Martinus, K. Tajudeen, K. Degila, K. Ogueji, K. Siminyu, J. Kreutzer, J. Webster, J. T. Ali, J. Abbott, I. Orife, I. Ezeani, I. A. Dangana, H. Kamper, H. Elsahar, G. Duru, G. Kioko, M. Espoir, E. van Biljon, D. Whitenack, C. Onyefuluchi, C. C. Emezue, B. F. P. Dossou, B. Sibanda, B. Bassey, A. Olabiyi, A. Ramkilowan, A. Oktem, A. Akinfaderin, and A. Bashir, “‘Participatory research for low-resourced machine translation: A case study in African languages,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020.
  40. 40.M. Artetxe, V. Goswami, S. Bhosale, A. Fan, and L. Zettlemoyer, “‘Revisiting machine translation for cross-lingual classification,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  41. 41.M. Abdin, J. Aneja, H. Behl, S. Bubeck, and Others, “‘Phi-4 technical report,” arXiv:2412.08905, 2024.
  42. 42.A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al., “Deepseek-v3 technical report,” arXiv:2412.19437, 2024.
  43. 43.Common Crawl. https://commoncrawl.org. Accessed: 2025-10-21.
  44. 44.S. Kudugunta, I. Caswell, B. Zhang, X. Garcia, D. Xin, A. Kusupati, R. Stella, A. Bapna, and O. Firat, “‘Madlad-400: A multilingual and document-level large audited dataset,” Advances in Neural Information Processing Systems, 2023.
  45. 45.NLLB Team, M. R. Costa-jussà, J. Cross, O. Çelebi, et al., “‘Scaling neural machine translation to 200 languages,” Nature, 2024.
  46. 46.T. Y. Zhou, Q. Xu, X. He, and T. Cohn, “‘Rethinking round-trip translation for machine translation evaluation,” in Findings of the Association for Computational Linguistics: EMNLP 2023, 2023.
  47. 47.H. Koretaka, T. Kajiwara, A. Fujita, and T. Ninomiya, “‘Mitigating domain mismatch in machine translation via paraphrasing,” in Proceedings of the 10th Workshop on Asian Translation, 2023.
  48. 48.D. Saunders and S. DeNeefe, “‘Domain adapted machine translation: What does catastrophic forgetting forget and why?,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024.
  49. 49.J. Shen, P.-J. Chen, M. Le, J. He, J. Gu, M. Ott, M. Auli, and M. Ranzato, “‘The source-target domain mismatch problem in machine translation,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics, 2021.
  50. 50.D. Saunders, “‘Domain adaptation and multi-domain adaptation for neural machine translation: A survey,” Journal of Artificial Intelligence Research, 2022.
  51. 51.M. Post, “‘A call for clarity in reporting BLEU scores,” in Proceedings of the Third Conference on Machine Translation: Research Papers, 2018.
  52. 52.M. Popović, “‘chrF++: words helping character n-grams,” in Proceedings of the Second Conference on Machine Translation, 2017.
  53. 53.S. Singh, F. Vargus, D. D’souza, B. F. Karlsson, A. Mahendiran, and Others, “‘Aya Dataset: an open-access collection for multilingual instruction tuning,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024.
  54. 54.I. Caswell, E. Nielsen, J. Luo, C. Cherry, G. Kovacs, H. Shemtov, P. Talukdar, D. Tewari, B. M. Diane, K. M. Doumbouya, et al., “‘SMOL: professionally translated parallel data for 115 underrepresented languages,” arXiv:2502.12301, 2025.
  55. 55.E. Bakouch, L. Ben Allal, A. Lozhkov, N. Tazi, L. Tunstall, C. M. Patiño, E. Beeching, A. Roucher, A. J. Reedi, Q. Gallouédec, K. Rasul, N. Habib, C. Fourrier, H. Kydlicek, G. Penedo, H. Larcher, M. Morlon, V. Srivastav, J. Lochner, X.-S. Nguyen, C. Raffel, L. von Werra, and T. Wolf, “‘SmolLM3: smol, multilingual, long-context reasoner,” Hugging Face Blog, 2025.
  56. 56.S. Huang, P. Li, Y. Hsu, K. Chen, Y. T. Lin, S. Hsiao, R. Tsai, and H. Lee, “‘Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 2024.
  57. 57.W. A. Gale and K. W. Church, “‘A program for aligning sentences in bilingual corpora,” Computational Linguistics, 1993.
  58. 58.R. C. Moore, “‘Fast and accurate sentence alignment of bilingual corpora,” in Proceedings of the 5th Conference of the Association for Machine Translation in the Americas (AMTA 2002), 2002.
  59. 59.B. Thompson and P. Koehn, “‘Vecalign: Improved sentence alignment in linear time and space,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019.
  60. 60.F. Feng, Y. Yang, D. Cer, N. Arivazhagan, and W. Wang, “‘Language-agnostic BERT sentence embedding,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022.
  61. 61.K. Heffernan, O. Çelebi, and H. Schwenk, “‘Bitext mining using distilled sentence representations for low-resource languages,” in Findings of the Association for Computational Linguistics: EMNLP 2022, 2022.
  62. 62.M. R. Costa-Jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, et al., “‘No language left behind: Scaling human-centered machine translation,” arXiv:2207.04672, 2022.
  63. 63.R. Sennrich, B. Haddow, and A. Birch, “‘Improving neural machine translation models with monolingual data,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics), 2016.
  64. 64.S. Edunov, M. Ott, M. Auli, and D. Grangier, “‘Understanding back-translation at scale,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018.
  65. 65.R. Haque, Y. Moslem, and A. Way, “‘Terminology-aware sentence mining for nmt domain adaptation: Adapt’s submission to the adap-mt 2020 english-to-hindi ai translation shared task,” in Proceedings of the 17th International Conference on Natural Language Processing (ICON): AdapMT 2020 Shared Task, 2020.
  66. 66.H. Schwenk, V. Chaudhary, S. Sun, H. Gong, and F. Guzmán, “‘WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 2021.
  67. 67.J. Tiedemann, “‘Parallel data, tools and interfaces in opus,” in Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC’12), 2012.
  68. 68.K. Nguyen and H. Daumé III, “‘Global Voices: Crossing borders in automatic news summarization,” in Proceedings of the 2nd Workshop on New Frontiers in Summarization, 2019.
  69. 69.A. Bapna, I. Caswell, J. Kreutzer, O. Firat, D. van Esch, A. Siddhant, M. Niu, P. N. Baljekar, X. Garcia, W. Macherey, T. Breiner, V. S. Axelrod, J. Riesa, Y. Cao, M. Chen, K. Macherey, M. Krikun, P. Wang, A. Gutkin, A. Shah, Y. Huang, Z. Chen, Y. Wu, and M. R. Hughes, “‘Building machine translation systems for the next thousand languages,” technical report, Google Research, 2022.
  70. 70.E. Nielsen, I. R. Caswell, J. Luo, and C. Cherry, “‘Alligators all around: Mitigating lexical confusion in low-resource machine translation,” in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), 2025.
  71. 71.L. Bandarkar, D. Liang, B. Muller, M. Artetxe, S. N. Shukla, D. Husa, N. Goyal, A. Krishnan, L. Zettlemoyer, and M. Khabsa, “‘The belebele benchmark: a parallel reading comprehension dataset in 122 language variants,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024.
  72. 72.P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord, “‘Think you have solved question answering? try arc, the ai2 reasoning challenge,” arXiv:1803.05457, 2018.
  73. 73.F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, et al., “‘Language models are multilingual chain-of-thought reasoners,” arXiv:2210.03057, 2022.
  74. 74.E. M. Ponti, G. Glavaš, O. Majewska, Q. Liu, I. Vulić, and A. Korhonen, “‘XCOPA: A multilingual dataset for causal commonsense reasoning,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020.
  75. 75.X. V. Lin, T. Mihaylov, M. Artetxe, T. Wang, S. Chen, D. Simig, M. Ott, N. Goyal, S. Bhosale, J. Du, et al., “‘Few-shot learning with multilingual language models,” arXiv:2112.10668, 2021.
  76. 76.Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al., “‘PIQA: reasoning about physical commonsense in natural language,” in Proceedings of the AAAI conference on artificial intelligence, 2020.
  77. 77.R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi, “‘HellaSwag: can a machine really finish your sentence?,” arXiv:1905.07830, 2019.
  78. 78.A. K. Upadhyay and H. K. Upadhya, “‘XNLI 2.0: improving xnli dataset and performance on cross lingual understanding (XLU),” in 2023 IEEE 8th International Conference for Convergence in Technology (I2CT), 2023.
  79. 79.A. Tikhonov and M. Ryabinin, “‘It’s all in the heads: Using attention heads as a baseline for cross-lingual transfer in commonsense reasoning,” arXiv:2106.12066, 2021.
  80. 80.S. Lin, J. Hilton, and O. Evans, “‘TruthfulQA: Measuring how models mimic human falsehoods,” arXiv:2109.07958, 2021.
  81. 81.L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou, “The language model evaluation harness,” Zenodo, 2024.
  82. 82.M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al., “‘Evaluating large language models trained on code,” arXiv:2107.03374, 2021.
  83. 83.M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, et al., “‘Challenging big-bench tasks and whether chain-of-thought can solve them,” arXiv:2210.09261, 2022.
  84. 84.D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman, “‘GPQA: a graduate-level google-proof q&a benchmark,” in First Conference on Language Modeling, 2024.
  85. 85.J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou, “‘Instructionfollowing evaluation for large language models,” arXiv:2311.07911, 2023.
  86. 86.D. S. Smart, “‘MultiWikiQA: a reading comprehension benchmark in 300+ languages,” arXiv:2509.04111v1, 2025.
  87. 87.E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “‘LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations, 2022.
  88. 88.A. Parrish, A. Chen, N. Nangia, V. Padmakumar, J. Phang, J. Thompson, P. M. Htut, and S. Bowman, “‘BBQ: A hand-built bias benchmark for question answering,” in Findings of the Association for Computational Linguistics: ACL 2022, 2022.
  89. 89.T. Hartvigsen, S. Gabriel, H. Palangi, M. Sap, D. Ray, and E. Kamar, “‘ToxiGen: A large-scale machine-generated dataset for implicit and adversarial hate speech detection,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 2022.
  90. 90.S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith, “‘RealToxicityPrompts: Evaluating neural toxic degeneration in language models,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020.
  91. 91.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “‘Attention is all you need,” in Advances in Neural Information Processing Systems, 2017.
  92. 92.R. Sennrich, B. Haddow, and A. Birch, “‘Neural machine translation of rare words with subword units,” in Proceedings of the 54th annual meeting of the association for computational linguistics, 2016.
  93. 93.D. P. Kingma and J. Ba, “‘Adam: A method for stochastic optimization,” in International Conference on Learning Representations, 2015.
  94. 94.N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “‘Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research, 2014.
  95. 95.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “‘Rethinking the inception architecture for computer vision,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  96. 96.E. Joanis, R. Knowles, R. Kuhn, S. Larkin, P. Littell, C.-k. Lo, D. Stewart, and J. Micher, “‘The Nunavut Hansard Inuktitut–English parallel corpus 3.0 with preliminary machine translation results,” in Proceedings of the Twelfth Language Resources and Evaluation Conference, 2020.
  97. 97.M. Post, T. Gowda, R. Grundkiewicz, H. Khayrallah, R. Jain, and M. Junczys-Dowmunt, “‘SOTASTREAM: A streaming approach to machine translation training,” arXiv:2308.07489, 2023.
  98. 98.L. Shen, W. Tan, S. Chen, Y. Chen, J. Zhang, H. Xu, B. Zheng, P. Koehn, and D. Khashabi, “‘The language barrier: Dissecting safety challenges of LLMs in multilingual contexts,” in Findings of the Association for Computational Linguistics: ACL 2024, 2024.
  99. 99.H. Zhao, C. Yuan, F. Huang, X. Hu, Y. Zhang, A. Yang, B. Yu, D. Liu, J. Zhou, J. Lin, et al., “‘Qwen3Guard technical report,” arXiv:2510.14276, 2025.
  100. 100.P. H. Martins, P. Fernandes, J. Alves, N. M. Guerreiro, R. Rei, D. M. Alves, J. Pombal, A. Farajian, M. Faysse, M. Klimaszewski, et al., “‘EuroLLM: Multilingual language models for europe,” Procedia Computer Science, 2025.
  101. 101.L. Dou, Q. Liu, F. Zhou, C. Chen, Z. Wang, Z. Jin, Z. Liu, T. Zhu, C. Du, P. Yang, et al., “‘Sailor2: Sailing in south-east asia with inclusive multilingual llms,” arXiv:2502.12982, 2025.
  102. 102.A. Omnilingual, G. Keren, A. Kozhevnikov, Y. Meng, C. Ropers, M. Setzler, S. Wang, I. Adebara, M. Auli, C. Balioglu, et al., “‘Omnilingual ASR: open-source multilingual speech recognition for 1600+ languages,” arXiv:2511.09690, 2025.
  103. 103.A. Romanou, N. Foroutan, A. Sotnikova, Z. Chen, S. H. Nelaturu, S. Singh, R. Maheshwary, M. Altomare, M. A. Haggag, A. Amayuelas, et al., “‘Include: Evaluating multilingual language understanding with regional knowledge,” arXiv:2411.19799, 2024.
  104. 104.H. Kydlíćek, G. Penedo, and L. von Werra, “‘FinePDFs.” https://huggingface.co/datasets/HuggingFaceFW/finepdfs, 2025. Hugging Face dataset.
  105. 105.O. De Gibert, G. Nail, N. Arefyev, M. Bañón, J. Van Der Linde, S. Ji, J. Zaragoza-Bernabeu, M. Aulamo, G. Ramírez-Sánchez, A. Kutuzov, et al., “‘A new massive multilingual dataset for high-performance language technologies,” arXiv:2403.14009, 2024.
  106. 106.T. Nguyen, C. V. Nguyen, V. D. Lai, H. Man, N. T. Ngo, F. Dernoncourt, R. A. Rossi, and T. H. Nguyen, “‘CulturaX: A cleaned, enormous, and multilingual dataset for large language models in 167 languages,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), 2024.
  107. 107.H. Laurençon, L. Saulnier, T. Wang, C. Akiki, A. Villanova del Moral, T. Le Scao, L. Von Werra, C. Mou, E. González Ponferrada, H. Nguyen, et al., “‘The bigscience roots corpus: A 1.6 tb composite multilingual dataset,” Advances in Neural Information Processing Systems, 2022.
  108. 108.G. Wenzek, M.-A. Lachaux, A. Conneau, V. Chaudhary, F. Guzmán, A. Joulin, and E. Grave, “‘CCNet: Extracting high quality monolingual datasets from web crawl data,” in Proceedings of the Twelfth Language Resources and Evaluation Conference, 2020.
  109. 109.L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “‘mT5: a massively multilingual pre-trained text-to-text transformer,” in Proceedings of NAACL, 2021.
  110. 110.L. Burchell, O. D. G. Bonet, N. Arefyev, M. Aulamo, M. Bañón, P. Chen, M. Fedorova, L. Guillou, B. Haddow, J. Hajic, et al., “‘An expanded massive multilingual dataset for high-performance language technologies (HPLT),” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025.
  111. 111.M. Aloui, H. Chouikhi, G. Chaabane, H. Kchaou, and C. Dhaouadi, “101 billion arabic words dataset,” arXiv:2405.01590, 2024.
  112. 112.M. S. U. R. Khan, P. Mehta, A. Sankar, U. Kumaravelan, S. Doddapaneni, S. Jain, A. Kunchukuttan, P. Kumar, R. Dabre, M. M. Khapra, et al., “IndicLLMSuite: a blueprint for creating pretraining and fine-tuning datasets for indian languages,” arXiv:2403.06350, 2024.
  113. 113.P. Shantipriya, L. Kusum, P. Shakshi, and M. Sanskruti, “‘Building pre-train llm dataset for the indic languages: A case study on hindi,” arXiv:2407.09855v1, 2024.
  114. 114.M. Faysse, P. Fernandes, N. M. Guerreiro, A. Loison, D. M. Alves, C. Corro, N. Boizard, J. Alves, R. Rei, P. H. Martins, et al., “‘CroissantLLM: a truly bilingual french-english language model,” arXiv:2402.00786, 2024.
  115. 115.M. Turker, M. E. Ari, and A. Han, “‘Vbart: The turkish llm,” arXiv:2403.01308, 2024.
  116. 116.X. Du, Z. Yu, S. Gao, D. Pan, Y. Cheng, Z. Ma, R. Yuan, X. Qu, J. Liu, T. Zheng, et al., “Chinese tiny llm: Pretraining a chinese-centric large language model,” arXiv:2404.04167, 2024.
  117. 117.V. Aryabumi, J. Dang, D. Talupuru, S. Dash, D. Cairuz, H. Lin, B. Venkitesh, M. Smith, J. A. Campos, Y. C. Tan, et al., “‘Aya 23: Open weight releases to further multilingual progress,” arXiv:2405.15032, 2024.
  118. 118.G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, et al., “‘Gemma 2: Improving open language models at a practical size,” arXiv:2408.00118, 2024.
  119. 119.H. W. Chung, N. Constant, X. Garcia, A. Roberts, Y. Tay, S. Narang, and O. Firat, “‘Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining,” arXiv:2304.09151, 2023.
  120. 120.M. Ali, M. Fromm, K. Thellmann, J. Ebert, A. A. Weber, R. Rutmann, C. Jain, M. Lübbering, D. Steinigen, J. Leveling, et al., “‘Teuken-7b-base & teuken-7b-instruct: Towards european llms,” arXiv:2410.03730, 2024.
  121. 121.N. Sengupta, S. K. Sahu, B. Jia, S. Katipomu, H. Li, F. Koto, W. Marshall, G. Gosal, C. Liu, Z. Chen, et al., “‘Jais and jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models,” arXiv:2308.16149, 2023.
  122. 122.M. Choudhury, S. Chauhan, R. J. Das, D. Sahnan, X. Han, et al., “‘Llama-3-Nanda-10B-Chat: An open generative large language model for hindi,” arXiv:2504.06011, 2025.
  123. 123.O. Gouvert, J. Hunter, J. Louradour, C. Cerisara, E. Dufraisse, Y. Sy, L. Rivière, J.-P. Lorré, and O.-F. community, “‘The lucie-7b llm and the lucie training dataset: Open resources for multilingual language generation,” arXiv:2503.12294, 2025.
  124. 124.R. Luukkonen, J. Burdge, E. Zosa, A. Talman, V. Komulainen, V. Hatanpää, P. Sarlin, and S. Pyysalo, “‘Poro 34b and the blessing of multilinguality,” arXiv:2404.01856, 2024.
  125. 125.Barcelona Supercomputing Center (Projecte AINA), “Àguila-7b: An open-source llm for catalan and spanish.” Project technical release / model documentation, 2023.
  126. 126.F. Koto, R. Joshi, N. Mukhituly, and Others, “‘Sherkala-Chat: Building a state-of-the-art llm for kazakh in a moderately resourced setting,” arXiv:2503.01493, 2025.
  127. 127.ISSAI, Nazarbayev University, “LLama-3.1-KazLLM-1.0-8B.” Model release, 2024.
  128. 128.D. Roussis, L. Voukoutis, G. Paraskevopoulos, S. Sofianopoulos, P. Prokopidis, V. Papavasileiou, A. Katsamanis, S. Piperidis, and V. Katsouros, “‘Krikri: Advancing open large language models for greek,” arXiv:2505.13772, 2025.
  129. 129.Jacaranda Health, “‘UlizaLlama: an open-access swahili large language model.” Model release / documentation, 2023.
  130. 130.M. Wortsman, G. Ilharco, S. Y. Gadre, R. Roelofs, R. Gontijo-Lopes, A. S. Morcos, H. Namkoong, A. Farhadi, Y. Carmon, S. Kornblith, and L. Schmidt, “‘Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time,” in Proceedings of the 39th International Conference on Machine Learning, 2022.
  131. 131.M. S. Matena and C. A. Raffel, “‘Merging models with fisher-weighted averaging,” Advances in Neural Information Processing Systems, 2022.
  132. 132.V. Gupta, S. Akle Serrano, and D. DeCoste, “‘Stochastic weight averaging in parallel: Large-batch training that generalizes well,” arXiv:2001.02312, 2020.
  133. 133.P. Yadav, D. Tam, L. Choshen, C. A. Raffel, and M. Bansal, “‘Ties-merging: Resolving interference when merging models,” Advances in Neural Information Processing Systems, 2023.
  134. 134.L. Yu, B. Yu, H. Yu, F. Huang, and Y. Li, “‘Language models are super mario: Absorbing abilities from homologous models as a free lunch,” in International Conference on Machine Learning, 2024.
  135. 135.H. A. A. K. Hammoud, U. Michieli, F. Pizzati, P. Torr, A. Bibi, B. Ghanem, and M. Ozay, “‘Model merging and safety alignment: One bad model spoils the bunch,” arXiv:2406.14563, 2024.
  136. 136.T. Wu, R. Yang, J. Li, P. Hu, N. Wong, and Y. Yang, “‘Shadow-FT: Tuning Instruct Model via Training on Paired Base Model,” arXiv:2505.12716, 2025.
  137. 137.A. Ahmadian, S. Goldfarb-Tarrant, B. Ermis, M. Fadaee, S. Hooker, et al., “‘Mix data or merge models? optimizing for diverse multi-task learning,” arXiv:2410.10801, 2024.
  138. 138.M. Tao, C. Zhang, Q. Huang, T. Ma, S. Huang, D. Zhao, and Y. Feng, “‘Unlocking the potential of model merging for low-resource languages,” in Findings of the Association for Computational Linguistics: EMNLP 2024, 2024.
  139. 139.B. Upadhayay and V. Behzadan, “‘TaCo: enhancing cross-lingual transfer for low-resource languages in LLMs through translation-assisted chain-of-thought processes,” in 5th Workshop on practical ML for limited/low resource settings, ICLR, 2024.
  140. 140.S. E. W. Contributors, “‘Simple english wikipedia.” https://simple.wikipedia.org/.
  141. 141.T. Community, “‘The tatoeba project.” https://tatoeba.org/, 2020.
  142. 142.X. Zhang, J. Zhao, and Y. LeCun, “‘Character-level convolutional networks for text classification,” in Advances in Neural Information Processing Systems, 2015.
  143. 143.S. Narayan, S. B. Cohen, and M. Lapata, “‘Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018.
  144. 144.M. Koupaee and W. Y. Wang, “‘WikiHow: A large scale text summarization dataset,” in Proceedings of the 2018 EMNLP Workshop on Analysis of Abusive Language, 2018.
  145. 145.D. Khashabi, Y. Kordi, T. Khot, A. Sabharwal, O. Tafjord, P. Clark, and H. Hajishirzi, “‘GooAQ: open question answering from questions asked to a large language model,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021.
  146. 146.T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. P. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Kelcey, J. Devlin, K. Lee, K. Toutanova, L. Jones, A. M. Dai, J. Uszkoreit, Q. V. Le, and S. Petrov, “‘Natural Questions: a benchmark for question answering research,” in Transactions of the Association for Computational Linguistics, 2019.
  147. 147.A. Fan, Y. Jernite, and J. Weston, “‘ELI5: Long form question answering,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  148. 148.P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “‘SQuAD: 100,000+ questions for machine comprehension of text,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2016.
  149. 149.M. Cettolo, C. Girardi, and M. Federico, “‘WIT3: web inventory of transcribed and translated talks,” in Proceedings of the 16th Annual Conference of the European Association for Machine Translation, 2012.
  150. 150.R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “‘Direct preference optimization: Your language model is secretly a reward model,” Advances in neural information processing systems, 2023.
  151. 151.Y. Dubois, C. X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. S. Liang, and T. B. Hashimoto, “‘AlpacaFarm: a simulation framework for methods that learn from human feedback,” Advances in Neural Information Processing Systems, 2023.

Citation

MLA
Zamir, S. W., et al. “BYOL: Bring Your Own Language Into LLMs”. arXiv, 2026, http://arxiv.org/abs/2601.10804v1.
APA
Zamir, S. W., Hamidouche, W., Amor, B. B., Marotti, L., Becker-Reshef, I., & Ferres, J. L. (2026). BYOL: Bring Your Own Language Into LLMs. arXiv. http://arxiv.org/abs/2601.10804v1
Chicago
Zamir, S. W., W. Hamidouche, B. B. Amor, L. Marotti, I. Becker-Reshef, and J. L. Ferres. 2026. “BYOL: Bring Your Own Language Into LLMs”. arXiv. http://arxiv.org/abs/2601.10804v1.
Harvard
Zamir, S.W. et al. (2026) “BYOL: Bring Your Own Language Into LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2601.10804v1.
Vancouver
1. Zamir SW, Hamidouche W, Amor BB, Marotti L, Becker-Reshef I, Ferres JL (2026) BYOL: Bring Your Own Language Into LLMs. arXiv

BibTeX

@article{zamir2026byol,
  title = {BYOL: Bring Your Own Language Into LLMs},
  author = {Zamir, Syed Waqas and Hamidouche, Wassim and Amor, Boulbaba Ben and Marotti, Luana and Becker-Reshef, Inbal and Ferres, Juan Lavista},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2601.10804v1},
  eprint = {2601.10804}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/