No Language Left Behind: Scaling Human-Centered Machine Translation

NLLB TeamMarta R. Costa-jussàJames CrossOnur ÇelebiMaha ElbayadKenneth HeafieldKevin HeffernanElahe KalbassiJanice LamDaniel Licht

article2022arXiv1,977 citations

Presents an open-source mixture-of-experts machine translation model capable of translating across more than 200 languages, delivering a 44% BLEU improvement over prior systems while establishing comprehensive benchmarks for low-resource evaluation and translation safety.

Listen

The vast majority of modern advances in artificial intelligence and machine translation have focused on a small subset of high-resource languages such as English, French, and Spanish, largely excluding the majority of the world’s languages. This technological imbalance severely limits digital inclusion, educational access, and cultural preservation for underserved, low-resource language communities globally. The article addresses this systemic inequity by developing an end-to-end, human-centered framework capable of delivering high-quality, safe, and fluent machine translation for more than 200 languages simultaneously.

The core objective was to design, train, and comprehensively evaluate a massively multilingual translation model that effectively doubles the language coverage of previous state-of-the-art systems while actively mitigating cross-language interference and minimizing translation toxicity.

To achieve this, the authors adopted a multi-stage approach combining human-centered qualitative research with large-scale computational engineering. They conducted in-depth interviews with 44 native speakers across 36 low-resource languages to establish foundational design principles and understand critical community needs. They constructed professionally translated seed and evaluation benchmarks, including Flores-200, which encompasses 3,001 human-translated sentences across 204 languages and establishes over 40,000 distinct translation directions. To address severe data scarcity, they implemented a lightweight fasttext-based language identification system capable of classifying over 200 languages and deployed teacher-student distillation models (LASER3) across 37.7 petabytes of web corpora to mine over 1.1 billion parallel sentence pairs. These mined and curated datasets were then utilized to train large-scale neural machine translation systems, most notably a 54.5-billion-parameter Sparsely Gated Mixture of Experts model utilizing conditional compute.

The evaluation revealed several key findings: First, the primary model achieved an overall 44% relative improvement in translation quality (BLEU score) over previous state-of-the-art benchmarks on Flores-101 while scaling to more than 200 languages. Second, the specialized sentence encoders substantially reduced multilingual alignment error rates—dropping average error on selected low-resource languages from 61% to under 1%—which directly enabled effective bitext mining even for languages with fewer than 100,000 native parallel sentences. Third, the conditional compute Mixture of Experts architecture successfully balanced cross-lingual knowledge transfer with minimal interference between unrelated languages, outperforming dense transformer baselines while preserving computational efficiency. Finally, combining automated filtering with newly curated toxic wordlists across 200 languages substantially mitigated the generation of harmful and hallucinatory translations.

These findings demonstrate that neural translation systems can successfully scale beyond 200 languages without sacrificing translation quality or safety. By significantly lowering the data threshold required to represent underserved languages, the work proves that high-quality translation can be democratized to support digital participation, educational equity, and knowledge platforms such as Wikipedia. However, because performance remains dependent on the volume of usable web text, extremely low-resource languages and predominantly oral languages still face notable challenges.

Organizations and researchers building upon this work should adopt the open-sourced models, benchmarks, and data pipelines while prioritizing human-in-the-loop validation for mission-critical deployments. Deploying smaller, distilled versions of the model (ranging from 600M to 3.3B parameters) is recommended for resource-constrained environments to balance inference cost with strong translation accuracy. Further research should focus on expanding support for oral and unstandardized languages, refining language identification for highly confusable dialects, and continually addressing cultural and domain generalization beyond web-centric corpora.

Cover for No Language Left Behind: Scaling Human-Centered Machine Translation

Abstract

Driven by the goal of eradicating language barriers on a global scale, machine translation has solidified itself as a key focus of artificial intelligence research today. However, such efforts have coalesced around a small subset of languages, leaving behind the vast majority of mostly low-resource languages. What does it take to break the 200 language barrier while ensuring safe, high quality results, all while keeping ethical considerations in mind? In No Language Left Behind, we took on this challenge by first contextualizing the need for low-resource language translation support through exploratory interviews with native speakers. Then, we created datasets and models aimed at narrowing the performance gap between low and high-resource languages. More specifically, we developed a conditional compute model based on Sparsely Gated Mixture of Experts that is trained on data obtained with novel and effective data mining techniques tailored for low-resource languages. We propose multiple architectural and training improvements to counteract overfitting while training on thousands of tasks. Critically, we evaluated the performance of over 40,000 different translation directions using a human-translated benchmark, Flores-200, and combined human evaluation with a novel toxicity benchmark covering all languages in Flores-200 to assess translation safety. Our model achieves an improvement of 44% BLEU relative to the previous state-of-the-art, laying important groundwork towards realizing a universal translation system. Finally, we open source all contributions described in this work, accessible at this https URL.

Table of Contents

  • 1. Introduction
  • 2. Human-Centered Low-Resource Language Translation
  • 2.1 Exploratory Interview Study Research Design
  • 2.1.1 Why should we prioritize low-resource languages?
  • 2.2 No Language Left Behind: Guiding Principles
  • 3. Languages
  • 4. Creating Professionally Translated Datasets: FLORES-200 and NLLB-Seed
  • 4.1 FLORES-200
  • 4.1.1 Benchmark Creation for Low-Resource Languages
  • 4.1.2 Benchmark Creation for Non-English Directions
  • 4.1.3 Flores-200 at a glance
  • 4.2 NLLB Seed Dataset
  • 4.3 NLLB Multi-Domain Dataset
  • 4.4 Conclusion
  • 5. Automatically Creating Translation Training Data for Hundreds of Languages
  • 5.1 Language Identification
  • 5.1.1 Related Work
  • 5.1.2 Models
  • 5.1.3 Improving LID with Linguistic Analysis
  • 5.1.4 Results
  • 5.2 Gathering and Cleaning Monolingual Data at Scale
  • 5.2.1 Description of our Monolingual Pipeline
  • 5.2.2 Monolingual Data at a Glance
  • 5.3 Mining Bitexts for Low-Resource Languages
  • 5.3.1 Related work
  • 5.3.2 Student-Teacher Mining Approach
  • 5.3.3 Language-Specific Encoder Training
  • 5.3.4 Mining at a Glance
  • 5.3.5 Limitations of Large-Scale Mining
  • 5.3.6 Ethical Considerations for Mining Research
  • 5.4 Conclusion
  • 6. Modeling
  • 6.1 Preliminaries
  • 6.1.1 Task Setup
  • 6.1.2 Ablation Dataset
  • 6.2 Conditional Compute for Massively Multilingual Machine Translation
  • 6.2.1 Vanilla Sparsely Gated MoE and its drawbacks for Low-Resource Languages
  • 6.2.2 Regularizing Massively Multilingual Mixtures of Experts
  • 6.2.3 Curriculum Learning
  • 6.2.4 Analysis of Multilingual Sparsely Gated MoE Models.
  • 6.3 Self-Supervision Strategies on Large-scale Monolingual Corpora
  • 6.3.1 Incorporating Self-Supervised Objectives with Multilingual Machine Translation
  • 6.3.2 Self-Supervised Learning Objectives
  • 6.3.3 Effect of Curriculum of Self-Supervision combined with Multilingual Machine Translation
  • 6.3.4 Effect of Self-Supervision Objectives
  • 6.3.5 Discussion
  • 6.4 Data Augmentation
  • 6.4.1 Different sources of data
  • 6.4.2 Data Tagging
  • 6.5 Bootstrapping models with NLLB-Seed
  • 6.5.1 Usefulness of NLLB-Seed
  • 6.5.2 Effect of NLLB-Seed on Backtranslation
  • 6.6 Human Evaluation
  • 6.7 Conclusion
  • 7. Evaluation
  • 7.1 Automatic Evaluation
  • 7.2 Human Evaluation
  • 7.2.1 Methodology
  • 7.2.2 Results
  • 7.3 Toxicity
  • 7.3.1 Preliminaries
  • 7.3.2 Toxicity Lists for 200 Languages
  • 7.3.3 Toxicity Detection
  • 7.3.4 Open Challenges in Toxicity for Machine Translation
  • 7.3.5 Ethical Considerations for Toxicity Research
  • 7.4 Conclusion
  • 8. Bringing it All Together
  • 8.1 Preparing the Data
  • 8.1.1 Training a Tokenizer for 200+ languages
  • 8.1.2 Datasets
  • 8.1.3 Large Scale Backtranslation
  • 8.1.4 Filtering Strategy
  • 8.1.5 Effect of using Different Data Sources on Performance
  • 8.1.6 The 200 Language Dataset
  • 8.2 Preparing the Model
  • 8.2.1 Does Self-supervised Learning help on top of Mining and Backtranslation?
  • 8.2.2 Scaling Model Architecture
  • 8.2.3 Designing an Optimized Training Curriculum
  • 8.2.4 The 200 Language Model: NLLB-200
  • 8.3 Results on Flores-200
  • 8.3.1 Performance on Flores-101 and Comparison to State-of-the-Art
  • 8.3.2 Performance on Flores-200
  • 8.3.3 Human Evaluation
  • 8.3.4 Prevalence of Toxicity
  • 8.4 Out-of-domain Generalization: Performance on non-Flores-200 Domains
  • 8.4.1 Public Benchmarks
  • 8.4.2 Results
  • 8.4.3 Effective Domain Adaptation with NLLB-MD
  • 8.5 Analysis of NLLB-200
  • 8.5.1 Language Co-location in NLLB-200 experts
  • 8.5.2 Effect of Phased Curriculum on Low-Resource Overfitting
  • 8.5.3 Impact of Multilingual Transfer
  • 8.6 Making Large Models More Accessible through Distillation
  • 8.6.1 Knowledge Distillation
  • 8.6.2 Creating Models Specialized for the Wikipedia Domain
  • 8.6.3 Distillation of NLLB-200, a 54B Parameter MoE Model
  • 8.6.4 Conclusion.
  • 8.7 Effectively Including Languages with Multiple Scripts and Related Languoids
  • 8.7.1 Transliteration
  • 8.7.2 Multidialectal Translation
  • 8.8 Environmental Impact of NLLB
  • 9. No Language Left Behind: Social Impact & Concluding Thoughts
  • 9.1 Expanding Information Access
  • 9.2 The Janus-faced Nature of Digital Participation
  • 9.3 The Future of NLLB: A Collective Responsibility
  • 10. Contributions
  • 11. Acknowledgements
  • References
  • Appendix A. Languages
  • A.1 Ethical Considerations around Language Standardization
  • Appendix B. Evaluation
  • B.1 Toxicity Lists
  • Appendix C. Data
  • Appendix D. Modeling
  • D.1 Ablation Dataset
  • D.2 SMT vs MMT
  • Appendix E. Bringing it All Together
  • E.1 Preparing the Data
  • E.1.1 Primary Dataset Composition
  • E.1.2 Importance of Backtranslation Quality on Model Scaling
  • E.1.3 Training Directions and Curriculum Buckets.
  • E.2 Results
  • E.2.1 Performance on African Languages
  • E.2.2 Comparison against Google Translate
  • E.3 Toxicity Evaluation
  • E.4 Out-of-domain Generalization: Performance on non Flores-200 Domains
  • E.5 Analysis of NLLB-200
  • E.6 Full Distillation Results.
  • Appendix F. Model Card - NLLB-200
  • Appendix G. Data Card for NLLB-Seed Data
  • Appendix H. Data Card for NLLB Multi-Domain Data
  • Appendix I. Data Card for Mined Bitext Metadata

Knowls

  1. Knowl 1 — NLLB-200 Sparse Mixture-of-Experts Translation Architecture

    model/method

    The NLLB-200 translation system is a massively multilingual sequence-to-sequence Transformer model scaled via sparsely gated Mixture-of-Experts (MoE) layers. The architecture contains 54.5 billion parameters across 24 encoder layers and 24 decoder layers with a model dimension of dmodel=2048d_{\text{model}} = 2048, a feed-forward network (FFN) dimension of dffn=8192d_{\text{ffn}} = 8192, and 16 attention heads. Computational cost per update (FLOPs) is kept equivalent to a 3.3 billion parameter dense Transformer.

    Every 4th Transformer block in both the encoder and decoder (fMoE=4f_{\text{MoE}} = 4) replaces the standard dense FFN with an MoE layer containing E=128E = 128 feed-forward expert subnetworks. For an input representation xt∈Rdmodelx_t \in \mathbb{R}^{d_{\text{model}}}, the gating network computes routing weights: Gt=Top-2-Gating(softmax(Wgxt))G_t = \text{Top-2-Gating}(\text{softmax}(W_g x_t)) where Wg∈RE×dmodelW_g \in \mathbb{R}^{E \times d_{\text{model}}}. The output of the MoE sublayer is the weighted sum of expert representations: MoE(xt)=∑e=1EGt,e⋅FFNe(xt)\text{MoE}(x_t) = \sum_{e=1}^E G_{t,e} \cdot \text{FFN}_e(x_t) where Gt,e>0G_{t,e} > 0 for at most k=2k = 2 experts, with no dispatch randomization. To balance expert allocation, the model is trained with an auxiliary load-balancing loss: LB=E⋅∑e=1Efepe,fe=1T∑t=1TI(argmax Gt=e),pe=1T∑t=1TGt,e\text{LB} = E \cdot \sum_{e=1}^E f_e p_e, \quad f_e = \frac{1}{T} \sum_{t=1}^T \mathbb{I}(\text{argmax } G_t = e), \quad p_e = \frac{1}{T} \sum_{t=1}^T G_{t,e} over mini-batches of TT tokens, weighted by 0.010.01 in the total training loss.

    The model employs Pre-LayerNorm, weight sharing across encoder input, decoder input, and decoder output embeddings, an overall dropout of 0.30.3, attention dropout of 0.10.1, and MoE Expert Output Masking with peom=0.2p_{\text{eom}} = 0.2. The vocabulary is tokenized using a joint SentencePiece model of 256,000 subwords trained on 100M sentences with temperature sampling T=5T = 5.

  2. Knowl 2 — FLORES-200 Multilingual Evaluation Benchmark

    definition

    FLORES-200 is a high-quality, professionally translated, many-to-many evaluation benchmark covering 204 languages (spanning 150 low-resource languages). It consists of 3,001 sentences sampled from English Wikimedia sources (one-third each from Wikinews, Wikijunior, and Wikivoyage across 842 articles with an average sentence length of 21 words). The dataset is divided into three fixed splits:

    • dev: 997 sentences (publicly available for validation and prompt tuning)
    • devtest: 1,012 sentences (publicly available for development evaluation)
    • test: 992 sentences (kept hidden behind an evaluation server for blinded benchmarking)

    Because the same 3,001 source sentences are translated into all 204 languages, the benchmark supports evaluation across 204×203=41,412204 \times 203 = 41,412 translation directions. The professional human translation workflow incorporates an initial language alignment phase, an initial 200-sentence translation phase with independent quality assurance (QA) arbitration, a full translation phase, and a final 20% sampled QA evaluation where language inclusion requires a minimum quality score threshold of 90%90\% (95%95\% for transliterated scripts).

  3. Knowl 3 — LASER3 Teacher-Student Multilingual Sentence Embedding Framework

    model/method

    LASER3 is a modular framework for creating massively multilingual sentence embedding spaces for bitext mining without retraining the global embedding model from scratch.

    A fixed multilingual teacher model, LASER2 (a 5-layer BiLSTM encoder trained on 93 languages), defines the anchor embedding space for English. For each new low-resource language (or cluster of related languages), a student model is instantiated as a 12-layer Transformer encoder (hidden dimension 1024, 4 attention heads, ∼250\sim 250M parameters) with its own dedicated SentencePiece vocabulary tailored to the student script.

    The student model is trained on available parallel text in the target language paired with English, complemented by 2 million sentences of English/Spanish bitext from CCMatrix to anchor the representations to the English embedding space. The training objective minimizes the cosine distance between the student representation of the non-English sentence and the teacher representation of the English sentence: Ldistill=1−cos⁡(Student(xtgt),Teacher(xsrc))\mathcal{L}_{\text{distill}} = 1 - \cos(\text{Student}(x_{\text{tgt}}), \text{Teacher}(x_{\text{src}})) When monolingual text is available, the student is jointly trained with a masked language modeling (MLM) loss on the target language text. LASER3 sentence encoders were trained for 148 languages, yielding an average error rate drop from 61%61\% (LASER) to 0.91%0.91\% (xsimx\text{sim}) on FLORES-200.

  4. Knowl 4 — Margin-Based Cross-Lingual Sentence Retrieval Score (xsim)

    equation

    Bitext mining and cross-lingual sentence retrieval accuracy are evaluated using a margin-based multilingual similarity score that penalizes sentences located in dense embedding regions. For a source sentence embedding xx and target sentence embedding yy, the margin score is defined as: score(x,y)=margin(cos⁡(x,y),∑z∈NNk(x)cos⁡(x,z)2k+∑v∈NNk(y)cos⁡(y,v)2k)\text{score}(x, y) = \text{margin}\left(\cos(x, y), \sum_{z \in \text{NN}_k(x)} \frac{\cos(x, z)}{2k} + \sum_{v \in \text{NN}_k(y)} \frac{\cos(y, v)}{2k}\right) where cos⁡(⋅,⋅)\cos(\cdot, \cdot) is the cosine similarity between sentence embeddings, NNk(u)\text{NN}_k(u) denotes the kk nearest neighbors of sentence embedding uu in the opposite language space, and k=4k = 4. Under the ratio margin function: margin(a,b)=ab\text{margin}(a, b) = \frac{a}{b} the score computes the ratio between the direct cosine similarity and the average cosine similarity of the local neighborhoods. In the bitext mining pipeline, candidate pairs (x,y)(x, y) are extracted as parallel data when score(x,y)≥1.06\text{score}(x, y) \ge 1.06.

  5. Knowl 5 — MoE Expert Output Masking (EOM) Regularization

    model/method

    MoE Expert Output Masking (EOM) is a regularization technique designed to mitigate overfitting in massively multilingual Sparsely Gated Mixture-of-Experts models, particularly on low-resource language pairs with small training volumes.

    In standard MoE layers with Top-kk gating (k=2k = 2), each input token xtx_t activates two expert subnetworks FFNe1(xt)\text{FFN}_{e_1}(x_t) and FFNe2(xt)\text{FFN}_{e_2}(x_t). Under EOM, the output of each selected expert is independently set to zero with probability peomp_{\text{eom}} before being combined with gating weights: FFN~e(xt)=mt,e⋅FFNe(xt),mt,e∼Bernoulli(1−peom)\widetilde{\text{FFN}}_{e}(x_t) = m_{t, e} \cdot \text{FFN}_{e}(x_t), \quad m_{t, e} \sim \text{Bernoulli}(1 - p_{\text{eom}}) EOM(xt)=∑e=1EGt,e⋅FFN~e(xt)\text{EOM}(x_t) = \sum_{e=1}^E G_{t,e} \cdot \widetilde{\text{FFN}}_{e}(x_t) where Gt,eG_{t,e} is the gating probability from the Top-2 routing network. Masking the outputs prior to combination strengthens the residual connection around the MoE block and prevents co-adaptation between expert subnetworks and subsequent layers. EOM leaves the token counts used in the load balancing loss LB\text{LB} unmasked, ensuring uniform expert utilization during training.

  6. Knowl 6 — Conditional MoE Routing (CMR) Layer

    model/method

    Conditional MoE Routing (CMR) is an architectural mechanism that allows a Transformer layer to dynamically allocate model capacity per token between a shared dense parameter block and an MoE block. A CMR layer consists of a shared dense feed-forward network FFNshared\text{FFN}_{\text{shared}} in parallel with an MoE block containing EE expert feed-forward subnetworks.

    For an input token representation xtx_t, a binary gate calculates an MoE routing probability: g(xt)=sigmoid(WCMR⋅xt)g(x_t) = \text{sigmoid}(W_{\text{CMR}} \cdot x_t) where WCMRW_{\text{CMR}} are learnable gating weights. The output of the CMR layer is computed as: CMR(xt)=(1−g(xt))⋅FFNshared(xt)+g(xt)⋅MoE(xt)\text{CMR}(x_t) = (1 - g(x_t)) \cdot \text{FFN}_{\text{shared}}(x_t) + g(x_t) \cdot \text{MoE}(x_t) To constrain capacity allocation, an auxiliary L1L_1 budget penalty is added to the training objective: LCMR=1T∑t=1T∣g(xt)−p∣L_{\text{CMR}} = \frac{1}{T} \sum_{t=1}^T |g(x_t) - p| where p∈[0,1]p \in [0, 1] is a target capacity budget hyperparameter (set empirically to p=0.8p = 0.8). In addition, a dropout fraction pcmrp_{\text{cmr}} of tokens per batch is randomly forced to bypass the MoE layer entirely (g(xt)=0g(x_t) = 0), directing them exclusively to FFNshared\text{FFN}_{\text{shared}}.

  7. Knowl 7 — Phased Curriculum Learning for Massively Multilingual Machine Translation

    algorithm

    Phased Curriculum Learning schedules the introduction of training language pairs to prevent low-resource directions from overfitting while allowing high-resource directions to train for hundreds of thousands of updates.

    Input: Total update steps TT, language pairs P\mathcal{P}, baseline unregularized run validation checkpoints
    Output: Trained multilingual translation model
    1. For each direction p∈Pp \in \mathcal{P}:
         Determine overfitting step sp=step of minimum validation perplexity in baselines_p = \text{step of minimum validation perplexity in baseline}
    2. Partition P\mathcal{P} into n=4n = 4 buckets {b0,b1,b2,b3}\{b_0, b_1, b_2, b_3\} based on median overfitting steps kik_i:
         k0=300kk_0 = 300\text{k} (pairs that converge slowly / high-resource)
         k1=130kk_1 = 130\text{k}
         k2=70kk_2 = 70\text{k}
         k3=30kk_3 = 30\text{k} (pairs that overfit rapidly / very low-resource)
    3. Initialize model parameters θ\theta at step t=0t = 0 with active dataset Dtrain=b0\mathcal{D}_{\text{train}} = b_0
    4. for update step t=1t = 1 to TT do:
         if t=T−k1t = T - k_1 (step 170k170\text{k}) then Dtrain←Dtrain∪b1\mathcal{D}_{\text{train}} \leftarrow \mathcal{D}_{\text{train}} \cup b_1
         if t=T−k2t = T - k_2 (step 230k230\text{k}) then Dtrain←Dtrain∪b2\mathcal{D}_{\text{train}} \leftarrow \mathcal{D}_{\text{train}} \cup b_2
         if t=T−k3t = T - k_3 (step 270k270\text{k}) then Dtrain←Dtrain∪b3\mathcal{D}_{\text{train}} \leftarrow \mathcal{D}_{\text{train}} \cup b_3
         Sample batch B∼DtrainB \sim \mathcal{D}_{\text{train}}
         Update θ\theta via Adam optimizer on cross-entropy and load-balancing losses
    5. return θ\theta

    On the 202-language dataset with T=300kT = 300\text{k}, 4-phase curriculum learning improves validation performance on low-resource and very low-resource into-English translation by +0.7+0.7 and +1.2+1.2 chrF++ compared to a naive count-based curriculum.

  8. Knowl 8 — Hybrid MMT and SMT Backtranslation with Finegrained Source Tagging

    model/method

    To generate diverse synthetic parallel data for underserved languages without compounding neural translation errors, backtranslation is generated from two complementary model families:

    1. Multilingual Neural Backtranslation (MmtBT): Generated using a dense 3.3B multilingual Transformer trained with a multitask objective combining translation (MMT) and denoising autoencoding (DAE) on monolingual text. MmtBT is applied to 261 English-centric directions where validation performance is below 30 spBLEU.
    2. Statistical Machine Translation Backtranslation (SmtBT): Generated using bilingual phrase-based statistical machine translation (Moses) models trained on primary and mined bitext. SmtBT is generated for 76 directions where neural baselines achieve <10<10 spBLEU out-of-English or <15<15 spBLEU into-English.

    To allow the translation model to distinguish data provenance and avoid overfitting on synthetic errors, every training example is prepended with a distinct provenance tag:

    • <PRIMARY_DATA> for human-translated and gold seed bitext
    • <MINED_DATA> for web-mined parallel sentences
    • <MMT_BT_DATA> for neural backtranslations
    • <SMT_BT_DATA> for phrase-based statistical backtranslations

    Finegrained tagging outperforms untagged training by +1.4+1.4 chrF++ across out-of-English directions and +1.1+1.1 chrF++ across into-English directions.

  9. Knowl 9 — Web-Scale Language Identification (LID-218) and Monolingual Data Filtering

    model/method

    Language identification for web crawling across 200+ languages is implemented using a linear fastText classifier trained on character nn-grams (n∈[2,5]n \in [2, 5]) with 256-dimensional embeddings, softmax cross-entropy loss, and temperature upsampling of training data with temperature 1/T=0.31/T = 0.3.

    To clean noisy web corpora (CommonCrawl and ParaCrawl, comprising 107.9 billion raw paragraphs), a multi-stage filtering pipeline is applied:

    1. Character Distribution Filtering: Sentences where <80%<80\% of characters fall outside the 95th percentile character distribution of the target language dev set are discarded. For large-alphabet languages (e.g., Japanese, Chinese), sentences containing <50%<50\% target Unicode script characters are discarded.
    2. English-Specific Filter: A dedicated binary fastText classifier discards spurious English sentences appearing within other language pages.
    3. Hierarchical Document/Sentence Consistency: Paragraph-level LID determines the language-specific sentence splitter; sentences whose sentence-level LID prediction disagrees with the paragraph-level LID prediction are removed.
    4. Score Thresholding: For languages exhibiting left-skewed LID probability distributions, a classifier threshold of 0.50.5 (or 0.90.9 for high-resource) is applied, while right-skewed low-resource languages use empirical peak thresholds between 0.20.2 and 0.40.4.
    5. Deduplication and Language Model Filtering: Deduplication is performed using hash normalization (xxh3_64), and KenLM perplexity filtering is applied to high-resource corpora to retain the top 30%30\% quality slice.
  10. Knowl 10 — Moderated Calibration Protocol for XSTS Human Translation Evaluation

    model/method

    Cross-lingual Semantic Text Similarity (XSTS) evaluates translation quality by scoring meaning preservation on a discrete 5-point scale from 1 (unrelated/incoherent) to 5 (completely semantically equivalent), prioritizing semantic adequacy over target fluency.

    To correct for annotator severity/leniency across disparate language pairs, all human evaluators rate a shared calibration set CsC_s of K=1000K = 1000 backtranslated FLORES-200 English sentences. For each translation direction ls→ltl_s \to l_t in study ss, the raw majority score Hls→lt(s)H^{(s)}_{l_s \to l_t} and the mean calibration score Cls→lt(s)C^{(s)}_{l_s \to l_t} are computed using sentence-level medians over evaluators. The calibrated score H~[mod]ls→lt(s)\widetilde{H}[\text{mod}]^{(s)}_{l_s \to l_t} is calculated via moderated calibration: C=Cls→lt(s)−Cˉ,Cˉ≈3.01C = C^{(s)}_{l_s \to l_t} - \bar{C}, \quad \bar{C} \approx 3.01 S=tanh⁡(−C)S = \tanh(-C) E={−tanh⁡(Hls→lt(s)−5)if C≤0tanh⁡(Hls→lt(s)−1)if C>0E = \begin{cases} -\tanh(H^{(s)}_{l_s \to l_t} - 5) & \text{if } C \le 0 \\ \tanh(H^{(s)}_{l_s \to l_t} - 1) & \text{if } C > 0 \end{cases} H~[mod]ls→lt(s)=Hls→lt(s)+S×E\widetilde{H}[\text{mod}]^{(s)}_{l_s \to l_t} = H^{(s)}_{l_s \to l_t} + S \times E

    Moderated calibration ensures scores remain bounded within [1,5][1, 5] while attenuating extreme shifts. Calibrated XSTS scores achieve higher Spearman correlation with automated metrics (R=0.710R = 0.710 with spBLEU, R=0.694R = 0.694 with sentence-level chrF++) than uncalibrated scores (R=0.625R = 0.625 with spBLEU).

  11. Knowl 11 — Toxicity-200 Lexicons and Imbalanced Bitext Filtering

    model/method

    Toxicity-200 provides human-translated and culturally adapted toxic wordlists across 200 languages (median list length: 143 entries; mean: 271 entries) covering profanities, slurs, sexual terms, hate speech, and bullying expressions, contextualized with parts-of-speech and disambiguating short nn-grams (0<n<40 < n < 4).

    Toxicity detectors identify toxic items via exact token matching (or SentencePiece matching for non-segmented languages such as Chinese, Japanese, Burmese, Assamese, and Khmer). In machine translation, the primary safety hazard is added toxicity (toxic items generated in the target translation that were absent in the source text), caused by hallucinations or crude lexical mistranslations.

    To prevent models from learning spurious associations that produce added toxicity, training bitext is filtered by toxicity imbalance: sentence pairs where the difference in the number of toxic items between source and target exceeds a threshold (2+ toxic items) are discarded. In bilingual models, toxicity filtering reduces added toxic items on FLORES-200 while simultaneously improving translation chrF++ (e.g., reducing toxic terms from 33 to 22 in eng-smo and from 14 to 0 in eng-umb).

  12. Knowl 12 — Empirical Translation Performance of NLLB-200 across Benchmarks

    empirical result

    The final 54.5B MoE model (NLLB-200) was evaluated across 40,602 translation directions in FLORES-200 and standard MT benchmarks:

    Benchmark / Subset eng_Latn-xx xx-eng_Latn
    all high low v.low all high low v.low
    FLORES-200 (chrF++) 45.3 54.9 41.9 39.5 56.8 63.5 54.4 54.4
    FLORES-200 (spBLEU) 27.1 38.3 23.1 21.3 38.0 44.7 35.5 35.6
    FLORES-101 (spBLEU) 34.0 - - - 41.2 - - -
    FLORES-101 DeltaLM Baseline 26.6 - - - 33.2 - - -

    On the FLORES-101 benchmark, NLLB-200 achieves an average of 24.024.0 spBLEU across all 10,000 directions, outperforming the previous state-of-the-art (DeltaLM at 16.716.7 spBLEU) by +7.3+7.3 spBLEU (44%44\% relative improvement). Across all 40,200 non-English directions (xx-yy) in FLORES-200, NLLB-200 attains 35.635.6 chrF++ (17.317.3 spBLEU) overall, with 35.435.4 chrF++ on 38,162 zero-shot directions compared to 39.739.7 chrF++ on supervised directions. In human evaluations using calibrated XSTS across 51 directions, NLLB-200 achieves an average score of 4.22/5.04.22 / 5.0, significantly outperforming a 3.3B dense baseline (3.66/5.03.66 / 5.0).

  13. Knowl 13 — Knowledge Distillation for Efficient Multilingual Model Deployment

    model/method

    To enable low-latency inference without requiring multi-GPU MoE routing, NLLB-200 is distilled into compact dense Transformer models:

    1. Full Multilingual Distillation: The 54.5B MoE NLLB-200 teacher is distilled into 1.3B and 615M dense student models supporting all 202 languages using online word-level knowledge distillation (minimizing cross-entropy against teacher soft probabilities LKD\mathcal{L}_{\text{KD}} over 200k updates). On FLORES-200 devtest, the 1.3B distilled model reaches 46.946.9 average chrF++ (+0.5+0.5 chrF++ over training a 1.3B dense model from scratch), and the 615M distilled model reaches 44.644.6 chrF++ (+0.3+0.3 chrF++ over the 615M baseline).

    2. Specialized Wikipedia Domain Distillation: For Wikipedia's Content Translation tool, a 1.3B dense teacher is fine-tuned on Wikipedia domain parallel text and distilled into a 500M parameter dense student covering 25 target languages and 74 translation directions. Offline sequence-level distillation (generating beam-search translations with beam size 4 on Wikipedia monolingual dumps) outperforms online distillation by +0.4+0.4 chrF++ on eng-xx and fra-xx, reaching 43.443.4 chrF++ on eng-xx (exceeding the 1.3B teacher by +0.4+0.4 chrF++).

  14. Knowl 14 — NLLB-Seed and NLLB Multi-Domain (NLLB-MD) Bootstrapping Datasets

    definition

    To bootstrap machine translation and evaluate out-of-domain transfer on low-resource languages, two professional human-translated corpora were created:

    1. NLLB-Seed: A parallel training corpus consisting of 6,193 English sentences sampled from Wikimedia's List of articles every Wikipedia should have (covering 11 topical categories) translated into 39 low-resource languages (and transliterated into 4 additional scripts, totaling 43 pairs). In bilingual baseline experiments, augmenting diminished public bitext with 6.2k NLLB-Seed sentences improves translation performance from 17.817.8 to 26.626.6 validation chrF++, which further rises to 30.130.1 chrF++ when combined with backtranslation.

    2. NLLB Multi-Domain (NLLB-MD): A multi-domain benchmark covering English paired with 6 languages (Central Aymara ayr_Latn, Bhojpuri bho_Deva, Dyula dyu_Latn, Friulian fur_Latn, Russian rus_Cyrl, and Wolof wol_Latn). It contains 3,000 professionally translated sentences in each of four distinct domains: News (WMT21 English-German dev set), Scripted Formal Speech (public spoken talks), Unscripted Informal Speech (multi-session conversational chat), and Health (WHO reports and TAUS COVID-19 data). Fine-tuning NLLB-200 without load balancing on 2,000 in-domain sentences improves out-of-domain accuracy across test sets by +7.7+7.7 chrF++ in Chat, +3.1+3.1 in News, +4.1+4.1 in Health, and +5.8+5.8 in Scripted Speech.

  15. Knowl 15 — Quantifying Dialectal Variation in Arabic Languoids via Dialectness Level

    equation

    Lexical distance between regional Arabic dialects (languoids) and Modern Standard Arabic (MSA, arb_Arab) is quantified using the Dialectness Level (DL) metric. DL represents the fraction of dialect tokens that do not appear in the vocabulary of normalized Modern Standard Arabic text.

    For a target languoid LL and reference Modern Standard Arabic corpus MM, the metrics are computed after Unicode character normalization (unifying forms of Alif, Hamzah, and numerals):

    • Corpus-level Dialectness Level (cDL): cDL(L,M)=∣{w∈VL∣w∉VM}∣∣VL∣c\text{DL}(L, M) = \frac{|\{w \in V_L \mid w \notin V_M\}|}{|V_L|} where VLV_L is the multiset of all tokens in the dialect evaluation corpus and VMV_M is the vocabulary set of MSA.
    • Sentence-level Dialectness Level (sDL): sDL(L,M)=1∣SL∣∑s∈SL∑w∈sI(w∉VM)∣s∣s\text{DL}(L, M) = \frac{1}{|S_L|} \sum_{s \in S_L} \frac{\sum_{w \in s} \mathbb{I}(w \notin V_M)}{|s|} where SLS_L is the set of sentences in the dialect corpus.

    On FLORES-200 devtest, lexical divergence against MSA ranges from Najdi (ars_Arab, sDL=3.01%s\text{DL} = 3.01\%, 96.596.5 chrF++) to South Levantine (ajp_Arab, sDL=42.13%s\text{DL} = 42.13\%, 47.347.3 chrF++) and Mesopotamian (acm_Arab, sDL=22.84%s\text{DL} = 22.84\%, 70.570.5 chrF++). Lower dialectness correlates directly with higher transfer performance from MSA.

Coverage note — Omitted general background reviews of earlier statistical/neural MT literature, qualitative societal interview narratives, and full language tables from the appendix, keeping all core technical algorithms, architectures, datasets, loss equations, and benchmark evaluation results.

References

  1. 1.Julien Abadji, Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. Towards a cleaner document-oriented multilingual crawled corpus. CoRR, abs/2201.06642, 2022. URL https://arxiv.org/abs/2201.06642.
  2. 2.Solomon Teferra Abate, Michael Melese, Martha Yifiru Tachbelie, Million Meshesha, Solomon Atinafu, Wondwossen Mulugeta, Yaregal Assabie, Hafte Abera, Binyam Ephrem Seyoum, Tewodros Abebe, et al. Parallel corpora for bi-directional statistical machine translation for seven ethiopian language pairs. In Proceedings of the First Workshop on Linguistic Resources for Natural Language Processing, pages 83–90, 2018.
  3. 3.Jade Abbott and Laura Martinus. Benchmarking neural machine translation for Southern African languages. In Proceedings of the 2019 Workshop on Widening NLP, pages 98–101, Florence, Italy, August 2019. Association for Computational Linguistics. URL https://aclanthology.org/W19-3632.
  4. 4.Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, and Stephan Vogel. The AMARA corpus: Building parallel language resources for the educational domain. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pages 1856–1862, Reykjavik, Iceland, May 2014. European Language Resources Association (ELRA). URL http://www.lrec-conf.org/proceedings/lrec2014/pdf/877_Paper.pdf.
  5. 5.Sadaf Abdul-Rauf and Holger Schwenk. On the Use of Comparable Corpora to Improve SMT performance. In EACL, pages 16–23, 2009. URL http://www.aclweb.org/anthology/E09-1003.
  6. 6.David Adelani, Dana Ruiter, Jesujoba Alabi, Damilola Adebonojo, Adesina Ayeni, Mofe Adeyemi, Ayodele Esther Awokoya, and Cristina España-Bonet. The effect of domain and diacritics in Yoruba–English neural machine translation. In Proceedings of the 18th Biennial Machine Translation Summit (Volume 1: Research Track), pages 61–75, Virtual, August 2021. Association for Machine Translation in the Americas. URL https://aclanthology.org/2021.mtsummit-research.6.
  7. 7.David Ifeoluwa Adelani, Jesujoba Oluwadara Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Chinenye Emezue, Colin Leong, Michael Beukman, Shamsuddeen Hassan Muhammad, Guyo Dub Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ayoade Ajibade, Tunde Oluwaseyi Ajayi, Yvonne Wambui Gitau, Jade Abbott, Mohamed Ahmed, Millicent Ochieng, Anuoluwapo Aremu, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Koffi Kalipe, Derguene Mbaye, Allahsera Auguste Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing Sibanda, Andiswa Bukula, and Sam Manthalu. A few thousand translations go a long way! leveraging pre-trained models for african news translation. CoRR, abs/2205.02022, 2022. doi: 10.48550/ARXIV.2205.02022. URL https://arxiv.org/abs/2205.02022.
  8. 8.Eneko Agirre, Daniel Cer, Mona Diab, and Aitor Gonzalez-Agirre. SemEval-2012 task 6: A pilot on semantic textual similarity. In **SEM 2012: The First Joint Conference on Lexical and Computational Semantics – Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation (SemEval 2012)*, pages 385–393, Montréal, Canada, June 2012. Association for Computational Linguistics. URL https://aclanthology.org/S12-1051.
  9. 9.Orevaoghene Ahia, Julia Kreutzer, and Sara Hooker. The low-resource double bind: An empirical study of pruning for low-resource machine translation. CoRR, abs/2110.03036, 2021. URL https://arxiv.org/abs/2110.03036.
  10. 10.Benjamin Akera, Jonathan Mukiibi, Lydia Sanyu Naggayi, Claire Babirye, Isaac Owomugisha, Solomon Nsumba, Joyce Nakatumba-Nabende, Engineer Bainomugisha, Ernest Mwebaze, and John Quinn. Machine translation for african languages: Community creation of datasets and models in uganda. In 3rd Workshop on African Natural Language Processing, 2022. URL https://openreview.net/forum?id=BK-z5qzEU-9.
  11. 11.Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ondřej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina España-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Kenneth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, Jungo Kasai, Daniel Khashabi, Kevin Knight, Tom Kocmi, Philipp Koehn, Nicholas Lourie, Christof Monz, Makoto Morishita, Masaaki Nagata, Ajay Nagesh, Toshiaki Nakazawa, Matteo Negri, Santanu Pal, Allahsera Auguste Tapo, Marco Turchi, Valentin Vydrin, and Marcos Zampieri. Findings of the 2021 conference on machine translation (WMT21). In Proceedings of the Sixth Conference on Machine Translation, pages 1–88, Online, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.wmt-1.1.
  12. 12.Amjad Almahairi, Nicolas Ballas, Tim Cooijmans, Yin Zheng, Hugo Larochelle, and Aaron Courville. Dynamic capacity networks. In Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML'16, page 2091 – 2100. JMLR.org, 2016.
  13. 13.Faisal Alshargi, Shahd Dibas, Sakhar Alkhereyf, Reem Faraj, Basmah Abdulkareem, Sane Yagi, Ouafaa Kacha, Nizar Habash, and Owen Rambow. Morphologically annotated corpora for seven Arabic dialects: Taizi, sanaani, najdi, jordanian, syrian, iraqi and Moroccan. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, 2019.
  14. 14.Charity Delmus Alupo, Daniel Omeiza, and David Vernon. Realizing the potential of ai in africa. Towards Trustworthy Artificial Intelligence Systems, 2021.
  15. 15.Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico, Christian Federmann, Dmitriy Genzel, Franscisco Guzmán, Junjie Hu, Macduff Hughes, Philipp Koehn, Rosie Lazar, Will Lewis, Graham Neubig, Mengmeng Niu, Alp Öktem, Eric Paquin, Grace Tang, and Sylwia Tur. TICO-19: the translation initiative for COvid-19. In Proceedings of the 1st Workshop on NLP for COVID-19 (Part 2) at EMNLP 2020, Online, December 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.nlpcovid19-2.5. URL https://aclanthology.org/2020.nlpcovid19-2.5.
  16. 16.Antonios Anastasopoulos, Ondřej Bojar, Jacob Bremerman, Roldano Cattoni, Maha Elbayad, Marcello Federico, Xutai Ma, Satoshi Nakamura, Matteo Negri, Jan Niehues, Juan Pino, Elizabeth Salesky, Sebastian Stüker, Katsuhito Sudoh, Marco Turchi, Alexander Waibel, Changhan Wang, and Matthew Wiesner. Findings of the IWSLT 2021 evaluation campaign. In Proceedings of the 18th International Conference on Spoken Language Translation (IWSLT 2021), pages 1–29, Bangkok, Thailand (online), August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.iwslt-1.1. URL https://aclanthology.org/2021.iwslt-1.1.
  17. 17.Patrick Andries. Proposition d'ajout de l'écriture tifinaghe. Organisation internationale de normalisation. Jeu universel des caractères codés sur octets (JUC). ORGANISATION INTERNATIONALE DE NORMALISATION, 2004.
  18. 18.Mohd Zeeshan Ansari, M. M. Sufyan Beg, Tanvir Ahmad, Mohd Jazib Khan, and Ghazali Wasim. Language identification of hindi-english tweets using code-mixed BERT. CoRR, abs/2107.01202, 2021. URL https://arxiv.org/abs/2107.01202.
  19. 19.Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George F. Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu. Massively multilingual neural machine translation in the wild: Findings and challenges. CoRR, abs/1907.05019, 2019. URL http://arxiv.org/abs/1907.05019.
  20. 20.Nigel Armstrong and Ian E. Mackenzie. Social levelling, or anti-standardization, pages 161–207. Palgrave Macmillan UK, London, 2013. ISBN 978-1-137-28439-6. doi: 10.1057/9781137284396_6. URL https://doi.org/10.1057/9781137284396_6.
  21. 21.Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al. Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information fusion, 58:82–115, 2020.
  22. 22.Mikel Artetxe and Holger Schwenk. Margin-based parallel corpus mining with multilingual sentence embeddings. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3197–3203, 2019a.
  23. 23.Mikel Artetxe and Holger Schwenk. Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond. TACL, pages 597–610, 2019b.
  24. 24.Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giri Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Mona T. Diab, Zornitsa Kozareva, and Ves Stoyanov. Efficient large scale language modeling with mixtures of experts. CoRR, abs/2112.10684, 2021. URL https://arxiv.org/abs/2112.10684.
  25. 25.Andoni Azpeitia, Thierry Etchegoyhen, and Eva Martínez Garcia. Weighted Set-Theoretic Alignment of Comparable Sentences. In BUCC, pages 41–45, 2017. URL http://aclweb.org/anthology/W17-2508.
  26. 26.Andoni Azpeitia, Thierry Etchegoyhen, and Eva Martínez Garcia. Extracting Parallel Sentences from Comparable Corpora with STACC Variants. In BUCC, May 2018.
  27. 27.Paul Azunre, Lawrence Adu-Gyamfi, Esther Appiah, Felix Akwerh, Salomey Osei, Cynthia Amoaba, Salomey Afua Addo, Edwin Buabeng-Munkoh, Nana Boateng, Franklin Adjei, and Bernard Adabankah. English-akuapem twi parallel corpus, January 2021a. URL https://doi.org/10.5281/zenodo.4432117.
  28. 28.Paul Azunre, Salomey Osei, Salomey Addo, Lawrence Asamoah Adu-Gyamfi, Stephen Moore, Bernard Adabankah, Bernard Opoku, Clara Asare-Nyarko, Samuel Nyarko, Cynthia Amoaba, et al. English-twi parallel corpus for machine translation. arXiv preprint arXiv:2103.15625, 2021b.
  29. 29.Paul Azunre, Salomey Osei, Salomey Addo, Lawrence Asamoah Adu-Gyamfi, Stephen Moore, Bernard Adabankah, Bernard Opoku, Clara Asare-Nyarko, Samuel Nyarko, Cynthia Amoaba, et al. NLP for ghanaian languages. arXiv preprint arXiv:2103.15475, 2021c.
  30. 30.Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. CoRR, abs/1607.06450, 2016. URL http://arxiv.org/abs/1607.06450.
  31. 31.Claire Babirye, Joyce Nakatumba-Nabende, Andrew Katumba, Ronald Ogwang, Jeremy Tusubira Francis, Jonathan Mukiibi, Medadi Ssentanda, Lilian D Wanzare, and Davis David. Building text and speech datasets for low resourced languages: A case of languages in east africa. In 3rd Workshop on African Natural Language Processing, 2022. URL https://openreview.net/forum?id=SO-U99z4U-q.
  32. 32.Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, ICLR 2015, 2015.
  33. 33.Loretta Baldassar, Mihaela Nedelcu, Laura Merla, and Raelene Wilding. Ict-based copresence in transnational families and communities: Challenging the premise of face-to-face proximity in sustaining relationships. Global Networks, 16(2):133–144, 2016.
  34. 34.Laith H. Baniata, Isaac. K. E. Ampomah, and Seyoung Park. A transformer-based neural machine translation model for arabic dialects that utilizes subword units. Sensors, 21 (19), 2021. ISSN 1424-8220. doi: 10.3390/s21196509. URL https://www.mdpi.com/1424-8220/21/19/6509.
  35. 35.Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, and Jaume Zaragoza. ParaCrawl: Web-scale acquisition of parallel corpora. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4555–4567, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.417. URL https://aclanthology.org/2020.acl-main.417.
  36. 36.Ankur Bapna, Isaac Caswell, Julia Kreutzer, Orhan Firat, Daan van Esch, Aditya Siddhant, Mengmeng Niu, Pallavi Baljekar, Xavier Garcia, Wolfgang Macherey, Theresa Breiner, Vera Axelrod, Jason Riesa, Yuan Cao, Mia Xu Chen, Klaus Macherey, Maxim Krikun, Pidong Wang, Alexander Gutkin, Apurva Shah, Yanping Huang, Zhifeng Chen, Yonghui Wu, and Macduff Hughes. Building machine translation systems for the next thousand languages, 2022. URL https://arxiv.org/abs/2205.03983.
  37. 37.Emily Bender. The #benderrule: On naming the languages we study and why it matters. The Gradient, 2019.
  38. 38.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pages 610–623, 2021.
  39. 39.Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In Proceedings of the 26th annual international conference on machine learning, pages 41–48, 2009.
  40. 40.Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville. Estimating or propagating gradients through stochastic neurons for conditional computation. CoRR, abs/1308.3432, 2013. URL http://arxiv.org/abs/1308.3432.
  41. 41.Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, and Rifat Shahriyar. Banglanlg: Benchmarks and resources for evaluating low-resource natural language generation in bangla. arXiv preprint arXiv:2205.11081, 2022.
  42. 42.Steven Bird. Designing for language revitalisation. In Gilles Adda, Khalid Choukri, Irmgarda Kasinskaite-Buddeberg, Joseph Mariani, Hélène Mazo, and Sakriani Sakti, editors, Language Technologies for All (LT4All), pages 296–299. European Language Resources Association (ELRA), 2019. URL https://en.unesco.org/LT4All. International Conference Language Technologies for All, LT4All ; Conference date: 04-12-2019 Through 06-12-2019.
  43. 43.Steven Bird and David Chiang. Machine translation for language preservation. In Proceedings of COLING 2012: Posters, pages 125–134, 2012.
  44. 44.Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. Language (technology) is power: A critical survey of “bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454–5476, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.485. URL https://aclanthology.org/2020.acl-main.485.
  45. 45.Su Lin Blodgett, Q Vera Liao, Alexandra Olteanu, Rada Mihalcea, Michael Muller, Morgan Klaus Scheuerman, Chenhao Tan, and Qian Yang. Responsible language technologies: Foreseeing and mitigating harms. In CHI Conference on Human Factors in Computing Systems Extended Abstracts, pages 1–3, 2022.
  46. 46.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146, 2017. ISSN 2307-387X.
  47. 47.Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. Findings of the 2016 conference on machine translation. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pages 131–198, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/W16-2301. URL https://aclanthology.org/W16-2301.
  48. 48.Marcel Bollmann. A large-scale comparison of historical text normalization systems. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3885–3898, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1389. URL https://aclanthology.org/N19-1389.
  49. 49.Houda Bouamor and Hassan Sajjad. H2@BUCC18: Parallel Sentence Extraction from Comparable Corpora Using Multilingual Sentence Embeddings. In BUCC, May 2018.
  50. 50.Houda Bouamor, Nizar Habash, Mohammad Salameh, Wajdi Zaghouani, Owen Rambow, Dana Abdulrahim, Ossama Obeid, Salam Khalifa, Fadhl Eryani, Alexander Erdmann, and Kemal Oflazer. The MADAR Arabic dialect corpus and lexicon. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan, May 2018. European Language Resources Association (ELRA). URL https://aclanthology.org/L18-1535.
  51. 51.Houda Bouamor, Sabit Hassan, and Nizar Habash. The MADAR shared task on Arabic fine-grained dialect identification. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pages 199–207, Florence, Italy, August 2019. Association for Computational Linguistics. doi: 10.18653/v1/W19-4622. URL https://aclanthology.org/W19-4622.
  52. 52.Pierre Bourdieu. Distinction: A social critique of the judgement of taste. Harvard University Press, 1987.
  53. 53.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. A large annotated corpus for learning natural language inference. In EMNLP, pages 632–642, Lisbon, Portugal, September 2015. Association for Computational Linguistics. doi: 10.18653/v1/D15-1075. URL https://aclanthology.org/D15-1075.
  54. 54.Peter F Brown, Vincent J Della Pietra, Stephen A Della Pietra, and Robert L Mercer. The mathematics of statistical machine translation: Parameter estimation. Computational linguistics, 1993.
  55. 55.Ralf D Brown. Non-linear mapping for improved identification of 1300+ languages. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 627–632, 2014.
  56. 56.Taina Bucher. Want to be on the top? algorithmic power and the threat of invisibility on facebook. New media & society, 14(7):1164–1180, 2012.
  57. 57.Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil. Model compression. In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 535–541, 2006.
  58. 58.Christian Buck and Philipp Koehn. Findings of the wmt 2016 bilingual document alignment shared task. In Proceedings of the First Conference on Machine Translation, pages 554–563, Berlin, Germany, August 2016. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/W/W16/W16-2347.
  59. 59.Lindsay Bywood, Panayota Georgakopoulou, and Thierry Etchegoyhen. Embracing the threat: machine translation as a solution for subtitling. Perspectives, 25(3):492–508, 2017.
  60. 60.Isaac Caswell, Ciprian Chelba, and David Grangier. Tagged back-translation. In Proceedings of the Fourth Conference on Machine Translation (Volume 1: Research Papers), pages 53–63, 2019.
  61. 61.Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus. In Proceedings of the 28th International Conference on Computational Linguistics, pages 6588–6608, Barcelona, Spain (Online), December 2020. International Committee on Computational Linguistics. doi: 10.18653/v1/2020.coling-main.579. URL https://aclanthology.org/2020.coling-main.579.
  62. 62.Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico. Report on the 11th IWSLT evaluation campaign. In Proceedings of the 11th International Workshop on Spoken Language Translation: Evaluation Campaign, pages 2–17, Lake Tahoe, California, December 2014. URL https://aclanthology.org/2014.iwslt-evaluation.1.
  63. 63.Mauro Cettolo, Marcello Federico, Luisa Bentivogli, Jan Niehues, Sebastian Stüker, Katsuhito Sudoh, Koichiro Yoshino, and Christian Federmann. Overview of the IWSLT 2017 evaluation campaign. In Proceedings of the 14th International Conference on Spoken Language Translation, pages 2–14, Tokyo, Japan, December 2017. International Workshop on Spoken Language Translation. URL https://aclanthology.org/2017.iwslt-1.1.
  64. 64.Guanhua Chen, Shuming Ma, Yun Chen, Dongdong Zhang, Jia Pan, Wenping Wang, and Furu Wei. Towards making the most of multilingual pretraining for zero-shot neural machine translation. CoRR, abs/2110.08547, 2021. URL https://arxiv.org/abs/2110.08547.
  65. 65.Zewen Chi, Li Dong, Shuming Ma, Shaohan Huang Xian-Ling Mao, Heyan Huang, and Furu Wei. mt6: Multilingual pretrained text-to-text transformer with translation pairs. arXiv preprint arXiv:2104.08692, 2021.
  66. 66.Kyunghyun Cho, Bart van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio. On the properties of neural machine translation: Encoder–decoder approaches. In Proceedings of SSST-8, Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation, pages 103–111, 2014.
  67. 67.Jack Choquette, Wishwesh Gandhi, Olivier Giroux, Nick Stam, and Ronny Krashinsky. Nvidia a100 tensor core gpu: Performance and innovation. IEEE Micro, 41(2):29–35, 2021.
  68. 68.Brian Christian. The alignment problem: Machine learning and human values. WW Norton & Company, 2020.
  69. 69.Christopher Cieri, Mike Maxwell, Stephanie Strassel, and Jennifer Tracey. Selection criteria for low resource language programs. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 4543–4549, 2016.
  70. 70.Donavyn Coffey. Māori are trying to save their language from big tech, April 2021. URL https://www.wired.co.uk/article/maori-language-tech.
  71. 71.Alexis Conneau and Guillaume Lample. Cross-lingual language model pretraining. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/file/c04c19c2c2474dbf5f7ac4372c5b9af1-Paper.pdf.
  72. 72.Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R Bowman, Holger Schwenk, and Veselin Stoyanov. Xnli: Evaluating cross-lingual sentence representations. arXiv preprint arXiv:1809.05053, 2018.
  73. 73.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 8440–8451. Association for Computational Linguistics, 2020. doi: 10.18653/v1/2020.acl-main.747. URL https://doi.org/10.18653/v1/2020.acl-main.747.
  74. 74.Marta R. Costa-jussà. An analysis of gender bias studies in natural language processing. Nature Machine Intelligence, 1(11):495–496, 2019.
  75. 75.Raj Dabre and Aneerav Sukhoo. Morisienmt: A dataset for mauritian creole machine translation. arXiv preprint arXiv:2206.02421, 2022.
  76. 76.Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully, Mitesh M. Khapra, and Pratyush Kumar. Indicbart: A pre-trained model for natural language generation of indic languages. CoRR, abs/2109.02903, 2021. URL https://arxiv.org/abs/2109.02903.
  77. 77.Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. Racial bias in hate speech and abusive language detection datasets. CoRR, abs/1905.12516, 2019. URL http://arxiv.org/abs/1905.12516.
  78. 78.Tullio De Mauro. Storia linguistica dell'Italia repubblicana: dal 1946 ai nostri giorni. Laterza, 2014. ISBN 9788858113622.
  79. 79.Kevin Degila, Godson Kalipe, Jamiil Touré Ali, and Momboladji Balogoun. Parallel text dataset for Neural Machine Translation (French -> Fongbe, French -> Ewe), November 2020. URL https://doi.org/10.5281/zenodo.4266935.
  80. 80.Stefano Demichelis and Jorgen W Weibull. Language, meaning, and games: A model of communication, coordination, and evolution. American Economic Review, 98(4):1292–1311, 2008.
  81. 81.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL, pages 4171–4186, 2019. URL https://aclanthology.org/N19-1423.
  82. 82.Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341, 2020.
  83. 83.Jesse Dodge, Taylor Prewitt, Remi Tachet Des Combes, Erika Odmark, Roy Schwartz, Emma Strubell, Alexandra Sasha Luccioni, Noah A Smith, Nicole DeCario, and Will Buchanan. Measuring the carbon intensity of ai in cloud instances. arXiv preprint arXiv:2206.05229, 2022.
  84. 84.Liam Donaldson and Paul Rutter. Healthier, fairer, safe: the global health journey 2007 – 2017. Technical report, World Health Organization, May 2017.
  85. 85.Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten Bosma, Zongwei Zhou, Tao Wang, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathy Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc V. Le, Yonghui Wu, Zhifeng Chen, and Claire Cui. Glam: Efficient scaling of language models with mixture-of-experts. CoRR, abs/2112.06905, 2021. URL https://arxiv.org/abs/2112.06905.
  86. 86.Jonathan Dunn. Mapping languages: The corpus of global language use. Language Resources and Evaluation, 54(4):999–1018, 2020.
  87. 87.Bernardt Duvenhage. Short text language identification for under resourced languages. CoRR, abs/1911.07555, 2019. URL http://arxiv.org/abs/1911.07555.
  88. 88.Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, and Katharina Kann. AmericasNLI: Evaluating zero-shot natural language understanding of pretrained multilingual models in truly low-resource languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6279–6299, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.435. URL https://aclanthology.org/2022.acl-long.435.
  89. 89.Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. Understanding back-translation at scale. In Proc. of EMNLP, 2018.
  90. 90.Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzman, and Philipp Koehn. A massive collection of cross-lingual web-document pairs. In EMNLP, pages 5960–5969, 2020.
  91. 91.Anca Elena-Bucea, Frederico Cruz-Jesus, Tiago Oliveira, and Pedro Simões Coelho. Assessing the role of age, education, gender and income on the digital divide: evidence for the european union. Information Systems Frontiers, 23(4):1007–1021, 2021.
  92. 92.Chris Chinenye Emezue and Bonaventure F. P. Dossou. MMTAfrica: Multilingual machine translation for African languages. In Proceedings of the Sixth Conference on Machine Translation, pages 398–411, Online, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.wmt-1.48.
  93. 93.Chris Chinenye Emezue and Femi Pancrace Bonaventure Dossou. Ffr v1. 1: Fon-french neural machine translation. In Proceedings of the The Fourth Widening Natural Language Processing Workshop, pages 83–87, 2020.
  94. 94.Cristina España-Bonet, Ádám Csaba Varga, Alberto Barrón-Cedeño, and Josef van Genabith. An Empirical Analysis of NMT-Derived Interlingual Embeddings and their Use in Parallel Sentence Identification. IEEE Journal of Selected Topics in Signal Processing, pages 1340–1348, 2017.
  95. 95.Thierry Etchegoyhen and Andoni Azpeitia. Set-Theoretic Alignment for Comparable Corpora. In ACL, pages 2009–2018, 2016. doi: 10.18653/v1/P16-1189. URL http://www.aclweb.org/anthology/P16-1189.
  96. 96.Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, and Armand Joulin. Beyond english-centric multilingual machine translation. The Journal of Machine Learning Research, 2020.
  97. 97.William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022. URL http://jmlr.org/papers/v23/21-0998.html.
  98. 98.Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. Language-agnostic bert sentence embedding, 2020. URL https://arxiv.org/abs/2007.01852.
  99. 99.Charles A Ferguson. Diglossia. word, 15(2):325–340, 1959.
  100. 100.Javier Ferrando, Gerard I. Gállego, Belen Alastruey, Carlos Escolano, and Marta R. Costa-jussà. Towards opening the black box of neural machine translation: Source and target interpretations of the transformer, 2022. URL https://arxiv.org/abs/2205.11631.
  101. 101.Markus Freitag, Yaser Al-Onaizan, and Baskaran Sankaran. Ensemble distillation for neural machine translation. CoRR, abs/1702.01802, 2017. URL http://arxiv.org/abs/1702.01802.
  102. 102.Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain. In Proceedings of the Sixth Conference on Machine Translation, pages 733–774, Online, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.wmt-1.73.
  103. 103.Batya Friedman and David G Hendry. Value sensitive design: Shaping technology with moral imagination. MIT Press, 2019.
  104. 104.Pascale Fung and Percy Cheung. Multi-level bootstrapping for extracting parallel sentences from a quasi-comparable corpus. In COLING 2004, 20th International Conference on Computational Linguistics, Proceedings of the Conference, 23-27 August 2004, Geneva, Switzerland, 2004. URL https://aclanthology.org/C04-1151/.
  105. 105.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. CoRR, abs/2009.11462, 2020. URL https://arxiv.org/abs/2009.11462.
  106. 106.Fantahun Gereme, William Zhu, Tewodros Ayall, and Dagmawi Alemu. Combating fake news in “low-resource” languages: Amharic fake news detection accompanied by resource crafting. Information, 12(1):20, 2021.
  107. 107.Jeff Good and Calvin Hendryx-Parker. Modeling contested categorization in linguistic databases. In Proceedings of the EMELD 2006 Workshop on Digital Language Documentation: Tools and standards: The state of the art, pages 20–22, 2006.
  108. 108.Google Jigsaw. Perpective api. https://www.perspectiveapi.com/, 2017. Accessed: 2022-05-03.
  109. 109.Mitchell A Gordon and Kevin Duh. Explaining sequence-level knowledge distillation as data-augmentation for neural machine translation. arXiv preprint arXiv:1912.03334, 2019.
  110. 110.Mitchell A. Gordon and Kevin Duh. Distill, adapt, distill: Training small, in-domain models for neural machine translation. CoRR, abs/2003.02877, 2020. URL https://arxiv.org/abs/2003.02877.
  111. 111.Cyril Goutte, Serge Léger, and Marine Carpuat. The NRC system for discriminating similar languages. In Proceedings of the first workshop on applying NLP tools to similar languages, varieties and dialects, pages 139–145, 2014.
  112. 112.Cyril Goutte, Serge Léger, Shervin Malmasi, and Marcos Zampieri. Discriminating similar languages: Evaluations and explorations. CoRR, abs/1610.00031, 2016. URL http://arxiv.org/abs/1610.00031.
  113. 113.Thamme Gowda, Zhao Zhang, Chris Mattmann, and Jonathan May. Many-to-English machine translation tools, data, and pretrained models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 306–316, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-demo.37. URL https://aclanthology.org/2021.acl-demo.37.
  114. 114.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc'Aurelio Ranzato, Francisco Guzmán, and Angela Fan. The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10:522–538, 2022. doi: 10.1162/tacl_a_00474. URL https://aclanthology.org/2022.tacl-1.30.
  115. 115.Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. Continuous measurement scales in human evaluation of machine translation. In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse, pages 33–41, Sofia, Bulgaria, August 2013. Association for Computational Linguistics. URL https://aclanthology.org/W13-2305.
  116. 116.Édouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomáš Mikolov. Learning word vectors for 157 languages. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), 2018.
  117. 117.Jiatao Gu, Hany Hassan, Jacob Devlin, and Victor O.K. Li. Universal neural machine translation for extremely low resource languages. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 344–354, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-1032. URL https://aclanthology.org/N18-1032.
  118. 118.Jiatao Gu, Yong Wang, Kyunghyun Cho, and Victor OK Li. Improved zero-shot neural machine translation via ignoring spurious correlations. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1258–1268, 2019.
  119. 119.Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge, Daniel Cer, Gustavo Hernandez Abrego, Keith Stevens, Noah Constant, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. Effective parallel corpus mining using bilingual sentence embeddings, 2018. URL https://arxiv.org/abs/1807.11906.
  120. 120.Udit Gupta, Mariam Elgamal, Gage Hills, Gu-Yeon Wei, Hsien-Hsin S Lee, David Brooks, and Carole-Jean Wu. Act: designing sustainable computer systems with an architectural carbon modeling tool. In Proceedings of the 49th Annual International Symposium on Computer Architecture, pages 784–799, 2022a.
  121. 121.Udit Gupta, Young Guen Kim, Sylvia Lee, Jordan Tse, Hsien-Hsin Sean Lee, Gu-Yeon Wei, David Brooks, and Carole-Jean Wu. Chasing carbon: The elusive environmental footprint of computing. IEEE Micro, 2022b.
  122. 122.Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc'Aurelio Ranzato. The FLORES evaluation datasets for low-resource machine translation: Nepali–English and Sinhala–English. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 6098–6111, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1632. URL https://aclanthology.org/D19-1632.
  123. 123.René Haas and Leon Derczynski. Discriminating between similar nordic languages. In Proceedings of the Eighth Workshop on NLP for Similar Languages, Varieties and Dialects, pages 67–75, Kiyv, Ukraine, April 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.vardial-1.8.
  124. 124.Nizar Habash. Introduction to arabic natural language processing. In Introduction to Arabic Natural Language Processing, 2010.
  125. 125.Nizar Habash, Ryan Roth, Owen Rambow, Ramy Eskander, and Nadi Tomeh. Morphological analysis and disambiguation for dialectal Arabic. In NAACL, pages 426–432, Atlanta, Georgia, June 2013. Association for Computational Linguistics. URL https://aclanthology.org/N13-1044.
  126. 126.Gilles Hacheme. English2gbe: A multilingual machine translation model for {Fon/Ewe}gbe. CoRR, abs/2112.11482, 2021. URL https://arxiv.org/abs/2112.11482.
  127. 127.Barry Haddow, Rachel Bawden, Antonio Valerio Miceli Barone, Jindřich Helcl, and Alexandra Birch. Survey of Low-Resource Machine Translation. Computational Linguistics, pages 1–67, 06 2022. ISSN 0891-2017. doi: 10.1162/coli_a_00446. URL https://doi.org/10.1162/coli_a_00446.
  128. 128.Asmelash Teka Hadgu, Gebrekirstos G. Gebremeskel, and Abel Aregawi. HornMT: Machine translation benchmark dataset for languages in the horn of africa. https://github.com/asmelashteka/HornMT, 2021.
  129. 129.Joan Kelly Hall. Teaching and researching: Language and culture. Routledge, 2013.
  130. 130.Harald Hammarström, Robert Forkel, Martin Haspelmath, and Sebastian Bank. Glottolog database 4.6, 2022.
  131. 131.Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, Shujie Liu, Tie-Yan Liu, Renqian Luo, Arul Menezes, Tao Qin, Frank Seide, Xu Tan, Fei Tian, Lijun Wu, Shuangzhi Wu, Yingce Xia, Dongdong Zhang, Zhirui Zhang, and Ming Zhou. Achieving human parity on automatic chinese to english news translation, 2018. URL https://arxiv.org/abs/1803.05567.
  132. 132.Einar Haugen. Planning for a standard language in modern norway. Anthropological Linguistics, 1(3):8–21, 1959. ISSN 00035483, 19446527. URL http://www.jstor.org/stable/30022188.
  133. 133.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In Proc. of CVPR, 2015.
  134. 134.Kenneth Heafield. KenLM: faster and smaller language model queries. In Proceedings of the EMNLP 2011 Sixth Workshop on Statistical Machine Translation, 2011.
  135. 135.Kevin Heffernan, Çelebi, and Holger Schwenk. Bitext mining using distilled sentence representations for low-resource languages, 2022. URL https://arxiv.org/abs/2205.12654.
  136. 136.Ulf Hermjakob, Jonathan May, and Kevin Knight. Out-of-the-box universal Romanization tool uroman. In Proceedings of ACL 2018, System Demonstrations, pages 13–18, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-4003. URL https://aclanthology.org/P18-4003.
  137. 137.Douglas Blanks Hindman. The rural-urban digital divide. Journalism & Mass Communication Quarterly, 77(3):549–560, 2000.
  138. 138.Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. URL https://arxiv.org/abs/1503.02531.
  139. 139.Phan Viet Hoang. Khmer natural language processing tookit. https://github.com/VietHoang1512/khmer-nltk, 2020.
  140. 140.Vu Cong Duy Hoang, Philipp Koehn, Gholamreza Haffari, and Trevor Cohn. Iterative back-translation for neural machine translation. In Proceedings of the 2nd Workshop on Neural Machine Translation and Generation, pages 18–24, 2018.
  141. 141.Md Zobaer Hossain, Md Ashraful Rahman, Md Saiful Islam, and Sudipta Kar. BanFakeNews: A dataset for detecting fake news in Bangla. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 2862–2871, Marseille, France, 2020. European Language Resources Association.
  142. 142.Catherine Howell and Darrell M. West. The internet as a human right, November 2016. URL https://www.brookings.edu/blog/techtank/2016/11/07/the-internet-as-a-human-right/.
  143. 143.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization. In ICML, pages 4411–4421, 2020.
  144. 144.Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, et al. Tutel: Adaptive mixture-of-experts at scale. arXiv preprint arXiv:2206.03382, 2022.
  145. 145.Tommi Jauhiainen, Krister Lindén, and Heidi Jauhiainen. Evaluation of language identification methods using 285 languages. In Proceedings of the 21st Nordic Conference on Computational Linguistics, pages 183–191, 2017.
  146. 146.Tommi Jauhiainen, Marco Lui, Marcos Zampieri, Timothy Baldwin, and Krister Lindén. Automatic language identification in texts: A survey. Journal of Artificial Intelligence Research, 65:675–782, 2019.
  147. 147.Isaac Johnson and Emily Lescak. Considerations for multilingual wikipedia research. arXiv preprint arXiv:2204.02483, 04 2022.
  148. 148.Melvin Johnson, Mike Schuster, Quoc V Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, et al. Google' s multilingual neural machine translation system: Enabling zero-shot translation. Transactions of the Association for Computational Linguistics, 5:339–351, 2017.
  149. 149.Pratik Joshi, Christain Barnes, Sebastin Santy, Simran Khanuja, Sanket Shah, Anirudh Srinivasan, Satwik Bhattamishra, Sunayana Sitaram, Monojit Choudhury, and Kalika Bali. Unsung challenges of building and deploying language technologies for low resource language communities. arXiv preprint arXiv:1912.03457, 2019.
  150. 150.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. The state and fate of linguistic diversity and inclusion in the nlp world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, 2020.
  151. 151.Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 427–431. Association for Computational Linguistics, April 2017.
  152. 152.Nal Kalchbrenner and Phil Blunsom. Recurrent continuous translation models. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1700–1709, Seattle, Washington, USA, October 2013. Association for Computational Linguistics. URL https://aclanthology.org/D13-1176.
  153. 153.Shivani Kapania, Oliver Siy, Gabe Clapper, Azhagu Meena SP, and Nithya Sambasivan. ” because ai is 100% right and safe” : User attitudes and sources of ai authority in india. In CHI Conference on Human Factors in Computing Systems, pages 1–18, 2022.
  154. 154.Alina Karakanta, Jon Dehdari, and Josef Genabith. Neural machine translation for low-resource languages without parallel corpora. Machine Translation, 32(1 – 2):167 – 189, jun 2018. ISSN 0922-6567. doi: 10.1007/s10590-017-9203-5. URL https://doi.org/10.1007/s10590-017-9203-5.
  155. 155.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pages 4171–4186, 2019.
  156. 156.Mahmoud Khonji, Youssef Iraqi, and Andrew Jones. Phishing detection: a literature survey. IEEE Communications Surveys & Tutorials, 15(4):2091–2121, 2013.
  157. 157.Elaine C Khoong and Jorge A Rodriguez. A research agenda for using machine translation in clinical medicine. Journal of General Internal Medicine, 37(5):1275–1277, 2022.
  158. 158.Yoon Kim and Alexander M Rush. Sequence-level knowledge distillation. In EMNLP, 2016.
  159. 159.Young Jin Kim, Ammar Ahmad Awan, Alexandre Muzio, Andrés Felipe Cruz-Salinas, Liyang Lu, Amr Hendy, Samyam Rajbhandari, Yuxiong He, and Hany Hassan Awadalla. Scalable and efficient moe training for multitask multilingual models. CoRR, abs/2109.10465, 2021. URL https://arxiv.org/abs/2109.10465.
  160. 160.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015. URL http://arxiv.org/abs/1412.6980.
  161. 161.Svetlana Kiritchenko, Isar Nejadgholi, and Kathleen C Fraser. Confronting abusive language online: A survey from the ethical and human rights perspective. Journal of Artificial Intelligence Research, 71:431–478, 2021.
  162. 162.Tom Kocmi, Dominik Macháček, and Ondřej Bojar. The Reality of Multi-Lingual Machine Translation. UFAL, Prague, Czechia, 2021.
  163. 163.Philipp Koehn. Europarl: A parallel corpus for statistical machine translation. In MT Summit, 2005.
  164. 164.Philipp Koehn. Statistical machine translation. Cambridge University Press, 2009.
  165. 165.Philipp Koehn and Ulrich Germann. The impact of machine translation quality on human post-editing. In Proceedings of the EACL 2014 Workshop on Humans and Computer-assisted Translation, pages 38–46, 2014.
  166. 166.Philipp Koehn and Rebecca Knowles. Six challenges for neural machine translation. In Proceedings of the First Workshop on Neural Machine Translation, pages 28–39, Vancouver, August 2017. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/W17-3204.
  167. 167.Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, et al. Moses: Open source toolkit for statistical machine translation. In Proceedings of the 45th annual meeting of the ACL on interactive poster and demonstration sessions, pages 177–180. Association for Computational Linguistics, 2007.
  168. 168.Philipp Koehn, Huda Khayrallah, Kenneth Heafield, and Mikel L. Forcada. Findings of the wmt 2018 shared task on parallel corpus filtering. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 726–739, Belgium, Brussels, October 2018. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/W18-6453.
  169. 169.Philipp Koehn, Francisco Guzmán, Vishrav Chaudhary, and Juan Pino. Findings of the WMT 2019 shared task on parallel corpus filtering for low-resource conditions. In Proceedings of the Fourth Conference on Machine Translation (Volume 3: Shared Task Papers, Day 2), 2019.
  170. 170.Philipp Koehn, Vishrav Chaudhary, Ahmed El-Kishky, Naman Goyal, Peng-Jen Chen, and Francisco Guzmán. Findings of the WMT 2020 shared task on parallel corpus filtering and alignment. In Proceedings of the Fifth Conference on Machine Translation, pages 726–742, Online, November 2020. Association for Computational Linguistics. URL https://aclanthology.org/2020.wmt-1.78.
  171. 171.Bruce Kogut and Anca Metiu. Open-source software development and distributed innovation. Oxford review of economic policy, 17(2):248–264, 2001.
  172. 172.Anastasia Kozyreva, Philipp Lorenz-Spreen, Ralph Hertwig, Stephan Lewandowsky, and Stefan M Herzog. Public attitudes towards algorithmic personalization and use of personal data online: Evidence from germany, great britain, and the united states. Humanities and Social Sciences Communications, 8(1):1–11, 2021.
  173. 173.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets. Transactions of the Association for Computational Linguistics, 10:50–72, January 2022. ISSN 2307-387X. doi: 10.1162/tacl_a_00447. URL https://doi.org/10.1162/tacl_a_00447.
  174. 174.Taku Kudo and John Richardson. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Eduardo Blanco and Wei Lu, editors, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018: System Demonstrations, Brussels, Belgium, October 31 - November 4, 2018, pages 66–71. Association for Computational Linguistics, 2018. doi: 10.18653/v1/d18-2012. URL https://doi.org/10.18653/v1/d18-2012.
  175. 175.Sneha Kudugunta, Yanping Huang, Ankur Bapna, Maxim Krikun, Dmitry Lepikhin, Minh-Thang Luong, and Orhan Firat. Beyond distillation: Task-level mixture-of-experts for efficient inference. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3577–3599, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp.304. URL https://aclanthology.org/2021.findings-emnlp.304.
  176. 176.Aman Kumar, Himani Shrotriya, Prachi Sahu, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Amogh Mishra, Mitesh M Khapra, and Pratyush Kumar. Indicnlg suite: Multilingual datasets for diverse nlg tasks in indic languages. arXiv preprint arXiv:2203.05437, 2022.
  177. 177.Sachin Kumar, Antonios Anastasopoulos, Shuly Wintner, and Yulia Tsvetkov. Machine translation into low-resource language varieties. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 110–121, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-short.16. URL https://aclanthology.org/2021.acl-short.16.
  178. 178.Anoop Kunchukuttan. The IndicNLP Library. https://github.com/anoopkunchukuttan/indic_nlp_library/blob/master/docs/indicnlp.pdf, 2020.
  179. 179.Keita Kurita, Anna Belova, and Antonios Anastasopoulos. Towards robust toxic content classification. CoRR, abs/1912.06872, 2019. URL http://arxiv.org/abs/1912.06872.
  180. 180.Remy Kusters, Dusan Misevic, Hugues Berry, Antoine Cully, Yann Le Cunff, Loic Dandoy, Natalia Díaz-Rodríguez, Marion Ficher, Jonathan Grizou, Alice Othmani, et al. Interdisciplinary research in artificial intelligence: Challenges and opportunities. Frontiers in Big Data, page 45, 2020.
  181. 181.Garry Kuwanto, Afra Feyza Akyürek, Isidora Chara Tourni, Siyang Li, and Derry Wijaya. Low-resource machine translation for low-resource languages: Leveraging comparable data, code-switching and compute resources. CoRR, abs/2103.13272, 2021. URL https://arxiv.org/abs/2103.13272.
  182. 182.Ivana Kvapilíková, Mikel Artetxe, Gorka Labaka amd Eneko Agirre, and Ondřej Bojar. Unsupervised multilingual sentence embeddings for parallel corpus mining. In ACL, 2020.
  183. 183.Niklas Laxström, Pau Giner, and Santhosh Thottingal. Content translation: Computer-assisted translation tool for wikipedia articles. CoRR, abs/1506.01914, 2015. URL http://arxiv.org/abs/1506.01914.
  184. 184.En-Shiun Lee, Sarubi Thillainathan, Shravan Nayak, Surangika Ranathunga, David Adelani, Ruisi Su, and Arya McCarthy. Pre-trained multilingual sequence-to-sequence models: A hope for low-resource language translation? In Findings of the Association for Computational Linguistics: ACL 2022, pages 58–67, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.6. URL https://aclanthology.org/2022.findings-acl.6.
  185. 185.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. CoRR, 2021. URL https://arxiv.org/abs/2107.06499.
  186. 186.Sangmin-Michelle Lee. The impact of using machine translation on efl students' writing. Computer Assisted Language Learning, 33(3):157–175, 2020.
  187. 187.Alyssa Lees, Daniel Borkan, Ian Kivlichan, Jorge Nario, and Tesh Goyal. Capturing covertly toxic speech via crowdsourcing. In Proceedings of the First Workshop on Bridging Human–Computer Interaction and Natural Language Processing, Online, April 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.hcinlp-1.3.
  188. 188.Heather Lent, Kelechi Ogueji, Miryam de Lhoneux, Orevaoghene Ahia, and Anders Søgaard. What a creole wants, what a creole needs. arXiv preprint arXiv:2206.00437, 2022.
  189. 189.Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. Gshard: Scaling giant models with conditional computation and automatic sharding. CoRR, abs/2006.16668, 2020. URL https://arxiv.org/abs/2006.16668.
  190. 190.Peggy Levitt and B Nadya Jaworsky. Transnational migration studies: Past developments and future trends. Annual review of sociology, 33:129, 2007.
  191. 191.Peggy Levitt and Deepak Lamba-Nieves. Social remittances revisited. Journal of ethnic and migration studies, 37(1):1–22, 2011.
  192. 192.Shahar Levy, Koren Lazar, and Gabriel Stanovsky. Collecting a large-scale gender bias dataset for coreference resolution and machine translation. arXiv preprint arXiv:2109.03858, 2021.
  193. 193.M. Paul Lewis, editor. Ethnologue: Languages of the World. SIL International, Dallas, TX, USA, sixteenth edition, 2009.
  194. 194.Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer. Base layers: Simplifying training of large, sparse models. In International Conference on Machine Learning, pages 6265–6274. PMLR, 2021.
  195. 195.Daniel Licht, Cynthia Gao, Janice Lam, Francisco Guzman, Mona Diab, and Philipp Koehn. Consistent human evaluation of machine translation across language pairs, 2022. URL https://arxiv.org/abs/2205.08533.
  196. 196.Pierre Lison and Jörg Tiedemann. Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles, 2016.
  197. 197.Rui Liu, Young Jin Kim, Alexandre Muzio, Barzan Mozafari, and Hany Hassan Awadalla. Gating dropout: Communication-efficient regularization for sparsely activated transformers. arXiv preprint arXiv:2205.14336, 2022.
  198. 198.Xuebo Liu, Longyue Wang, Derek F. Wong, Liang Ding, Lidia S. Chao, Shuming Shi, and Zhaopeng Tu. On the complementarity between pre-training and back-translation for neural machine translation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 2900–2907, Punta Cana, Dominican Republic, November 2021a. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp.247. URL https://aclanthology.org/2021.findings-emnlp.247.
  199. 199.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742, 2020.
  200. 200.Zihan Liu, Genta Indra Winata, and Pascale Fung. Continual mixed-language pre-training for extremely low-resource neural machine translation. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 2706–2718, Online, August 2021b. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-acl.239. URL https://aclanthology.org/2021.findings-acl.239.
  201. 201.Adam Lopez. Statistical machine translation. ACM Computing Surveys (CSUR), 40(3): 1–49, 2008.
  202. 202.Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 12-in-1: Multi-task vision and language representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10437–10446, 2020.
  203. 203.Stefano Lusito, Edoardo Ferrante, and Jean Maillard. Text normalization for endangered languages: the case of Ligurian. arXiv preprint arXiv:2206.07861, 2022.
  204. 204.Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, Alexandre Muzio, Saksham Singhal, Hany Hassan Awadalla, Xia Song, and Furu Wei. Deltalm: Encoder-decoder pre-training for language generation and translation by augmenting pretrained multilingual encoders. CoRR, abs/2106.13736, 2021. URL https://arxiv.org/abs/2106.13736.
  205. 205.Sean MacAvaney, Hao-Ren Yao, Eugene Yang, Katina Russell, Nazli Goharian, and Ophir Frieder. Hate speech detection: Challenges and solutions. PloS one, 14(8):e0221152, 2019.
  206. 206.Yash Madhani, Sushane Parthan, Priyanka Bedekar, Ruchi Khapra, Vivek Seshadri, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh M Khapra. Aksharantar: Towards building open transliteration tools for the next billion users. arXiv preprint arXiv:2205.03018, 2022.
  207. 207.Manuel Mager, Arturo Oncevay, Abteen Ebrahimi, John Ortega, Annette Rios, Angela Fan, Ximena Gutierrez-Vasques, Luis Chiruzzo, Gustavo Giménez-Lugo, Ricardo Ramos, Ivan Vladimir Meza Ruiz, Rolando Coto-Solano, Alexis Palmer, Elisabeth Mager-Hois, Vishrav Chaudhary, Graham Neubig, Ngoc Thang Vu, and Katharina Kann. Findings of the AmericasNLP 2021 shared task on open machine translation for indigenous languages of the Americas. In Proceedings of the First Workshop on Natural Language Processing for Indigenous Languages of the Americas, pages 202–217, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.americasnlp-1.23. URL https://aclanthology.org/2021.americasnlp-1.23.
  208. 208.Alexandre Magueresse, Vincent Carles, and Evan Heetderks. Low-resource languages: A review of past work and future challenges. arXiv preprint arXiv:2006.07264, 2020.
  209. 209.Laurette Marais, Ilana Wilken, Nina Van Niekerk, and Karen Calteaux. Mburisano covid-19 multilingual corpus. https://hdl.handle.net/20.500.12185/536, 2021.
  210. 210.Benjamin Marie, Raphael Rubino, and Atsushi Fujita. Tagged back-translation revisited: Why does it really work? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5990–5997, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.532. URL https://aclanthology.org/2020.acl-main.532.
  211. 211.Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, and David Yarowsky. The Johns Hopkins University Bible corpus: 1600+ tongues for typological exploration. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 2884–2892, Marseille, France, May 2020. European Language Resources Association. ISBN 979-10-95546-34-4. URL https://aclanthology.org/2020.lrec-1.352.
  212. 212.Leland McInnes, John Healy, Nathaniel Saul, and Lukas Grossberger. Umap: Uniform manifold approximation and projection. The Journal of Open Source Software, 3(29):861, 2018.
  213. 213.Cindy A. McKellar. Autshumato machine translation evaluation set. In Centre for Text Technology (CTexT), 2017.
  214. 214.Sharon J McLennan. Techno-optimism or information imperialism: Paradoxes in online networking, social media and development. Information Technology for Development, 22 (3):380–399, 2016.
  215. 215.Paul McNamee. Language identification: a solved problem suitable for undergraduate instruction. Journal of computing sciences in colleges, 20(3):94–101, 2005.
  216. 216.Sabrina J Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gallé, Arun Raja, Chenglei Si, Wilson Y Lee, Benoît Sagot, et al. Between words and characters: A brief history of open-vocabulary modeling and tokenization in nlp. arXiv preprint arXiv:2112.10508, 2021.
  217. 217.Tomas Mikolov and Geoffrey Zweig. Context dependent recurrent neural network language model. In 2012 IEEE Spoken Language Technology Workshop (SLT), pages 234–239. IEEE, 2012.
  218. 218.Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr, et al. A large-scale study of machine translation in the turkic languages. arXiv preprint arXiv:2109.04593, 2021.
  219. 219.Pushkar Mishra, Helen Yannakoudakis, and Ekaterina Shutova. Tackling online abuse: A survey of automated abuse detection methods. CoRR, abs/1908.06024, 2019. URL http://arxiv.org/abs/1908.06024.
  220. 220.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT '19*, page 220 – 229, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450361255. doi: 10.1145/3287560.3287596. URL https://doi.org/10.1145/3287560.3287596.
  221. 221.Aaron Mueller, Garrett Nicolai, Arya D McCarthy, Dylan Lewis, Winston Wu, and David Yarowsky. An analysis of massively multilingual neural machine translation for low-resource languages. In Proceedings of The 12th language resources and evaluation conference, pages 3710–3718, 2020.
  222. 222.Namrata Mukhija, Monojit Choudhury, and Kalika Bali. Designing language technologies for social good: The road not taken. arXiv preprint arXiv:2110.07444, 2021.
  223. 223.Dragos Stefan Munteanu and Daniel Marcu. Improving Machine Translation Performance by Exploiting Non-Parallel Corpora. Computational Linguistics, 31(4):477–504, 2005. URL http://www.aclweb.org/anthology/J05-4003.
  224. 224.Toshiaki Nakazawa, Chenchen Ding, Raj Dabre, Anoop Kunchukuttan, Nobushige Doi, Yusuke Oda, Ondřej Bojar, Shantipriya Parida, Isao Goto, and Hidaya Mino, editors. Proceedings of the 6th Workshop on Asian Translation, Hong Kong, China, November 2019. Association for Computational Linguistics. URL https://aclanthology.org/D19-5200.
  225. 225.Toshiaki Nakazawa, Hideki Nakayama, Chenchen Ding, Raj Dabre, Shohei Higashiyama, Hideya Mino, Isao Goto, Win Pa Pa, Anoop Kunchukuttan, Shantipriya Parida, Ondřej Bojar, Chenhui Chu, Akiko Eriguchi, Kaori Abe, Yusuke Oda, and Sadao Kurohashi. Overview of the 8th workshop on Asian translation. In Proceedings of the 8th Workshop on Asian Translation (WAT2021), pages 1–45, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.wat-1.1. URL https://aclanthology.org/2021.wat-1.1.
  226. 226.Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Meressa Berhe, Mofetoluwa Adeyemi, Masabata Mokgesi-Selinga, Lawrence Okegbemi, Laura Martinus, Kolawole Tajudeen, Kevin Degila, Kelechi Ogueji, Kathleen Siminyu, Julia Kreutzer, Jason Webster, Jamiil Toure Ali, Jade Abbott, Iroro Orife, Ignatius Ezeani, Idris Abdulkadir Dangana, Herman Kamper, Hady Elsahar, Goodness Duru, Ghollah Kioko, Murhabazi Espoir, Elan van Biljon, Daniel Whitenack, Christopher Onyefuluchi, Chris Chinenye Emezue, Bonaventure F. P. Dossou, Blessing Sibanda, Blessing Bassey, Ayodele Olabiyi, Arshath Ramkilowan, Alp Öktem, Adewale Akinfaderin, and Abdallah Bashir. Participatory research for low-resourced machine translation: A case study in African languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2144–2160, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.findings-emnlp.195. URL https://aclanthology.org/2020.findings-emnlp.195.
  227. 227.Toan Q. Nguyen and David Chiang. Transfer learning across low-resource, related languages for neural machine translation. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 296–301, Taipei, Taiwan, November 2017. Asian Federation of Natural Language Processing. URL https://aclanthology.org/I17-2050.
  228. 228.Abu Sadat Nurullah. Globalisation as a challenge to islamic cultural identity. International Journal of Interdisciplinary Social Sciences, 3(6):45–52, 2008.
  229. 229.Akintunde Oladipo, Odunayo Ogundepo, Kelechi Ogueji, and Jimmy Lin. An exploration of vocabulary size and transfer effects in multilingual language models for african languages. In 3rd Workshop on African Natural Language Processing, 2022. URL https://openreview.net/forum?id=HOZmF9MV8Wc.
  230. 230.Iroro Orife, Julia Kreutzer, Blessing Sibanda, Daniel Whitenack, Kathleen Siminyu, Laura Martinus, Jamiil Toure Ali, Jade Z. Abbott, Vukosi Marivate, Salomon Kabongo, Musie Meressa, Espoir Murhabazi, Orevaoghene Ahia, Elan Van Biljon, Arshath Ramkilowan, Adewale Akinfaderin, Alp Öktem, Wole Akin, Ghollah Kioko, Kevin Degila, Herman Kamper, Bonaventure Dossou, Chris Emezue, Kelechi Ogueji, and Abdallah Bashir. Masakhane - machine translation for africa. CoRR, abs/2003.11529, 2020. URL https://arxiv.org/abs/2003.11529.
  231. 231.Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary. Asynchronous Pipeline for Processing Huge Corpora on Medium to Low Resource Infrastructures. In Piotr Bański, Adrien Barbaresi, Hanno Biber, Evelyn Breiteneder, Simon Clematide, Marc Kupietz, Harald Lüngen, and Caroline Iliadi, editors, 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7), pages 9 – 16, Cardiff, United Kingdom, July 2019. Leibniz-Institut für Deutsche Sprache. doi: 10.14618/IDS-PUB-9021. URL https://hal.inria.fr/hal-02148693.
  232. 232.Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-4009. URL https://aclanthology.org/N19-4009.
  233. 233.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318, 2002.
  234. 234.David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021.
  235. 235.Amandalynne Paullada, Inioluwa Deborah Raji, Emily M Bender, Emily Denton, and Alex Hanna. Data and its (dis) contents: A survey of dataset development and use in machine learning research. Patterns, 2(11):100336, 2021.
  236. 236.Barbara Plank. What to do about non-standard (or non-canonical) language in NLP. In Proceedings of the 13th Conference on Natural Language Processing (KONVENS 2016), pages 13–20, 2016. doi: 10.18653/v1/D16-1163. URL https://aclanthology.org/D16-1163.
  237. 237.Maja Popović. chrf++: words helping character n-grams. In Proceedings of the Second Conference on Machine Translation, Volume 2: Shared Task Papers, pages 612–618, Copenhagen, Denmark, September 2017. Association for Computational Linguistics. URL http://www.aclweb.org/anthology/W17-4770.
  238. 238.Matt Post. A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186–191, Brussels, Belgium, October 2018. Association for Computational Linguistics. doi: 10.18653/v1/W18-6319. URL https://aclanthology.org/W18-6319.
  239. 239.Manasa Prasad, Theresa Breiner, and Daan van Esch. Mining training data for language modeling across the world’s languages. In SLTU, pages 61–65, 2018.
  240. 240.Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. BPE-dropout: Simple and effective subword regularization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1882–1892, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.170. URL https://www.aclweb.org/anthology/2020.acl-main.170.
  241. 241.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8): 9, 2019. URL https://d4mucfpksywv.cloudfront.net/better-language-models/language-models.pdf.
  242. 242.Alexandre Rafalovitch and Robert Dale. United Nations general assembly resolutions: A six-language parallel corpus. In Proceedings of the MT Summit XII, pages 292–299, Ottawa, Canada, 2014.
  243. 243.Jenalea Rajab. Effect of tokenisation strategies for low-resourced southern african languages. In 3rd Workshop on African Natural Language Processing, 2022. URL https://openreview.net/forum?id=SpMeq5M48W9.
  244. 244.Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He. Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale. arXiv preprint arXiv:2201.05596, 2022.
  245. 245.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100, 000+ questions for machine comprehension of text. In EMNLP, 2016.
  246. 246.Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kakwani, Navneet Kumar, Aswin Pradeep, Srihari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Shantadevi Khapr. Samanantar: The largest publicly available parallel corpora collection for 11 indic languages. Transactions of the Association for Computational Linguistics, 10:145–162, 2022. doi: 10.1162/tacl_a_00452. URL https://aclanthology.org/2022.tacl-1.9.
  247. 247.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. Comet: A neural framework for mt evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, 2020.
  248. 248.Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence embeddings using siamese bert-networks. In EMNLP, pages 3982–3992, 2019.
  249. 249.Nils Reimers and Iryna Gurevych. Making monolingual sentence embeddings multilingual using knowledge distillation. In EMNLP, pages 4512–4525, 2020.
  250. 250.Shuo Ren, Zhirui Zhang, Shujie Liu, Ming Zhou, and Shuai Ma. Unsupervised neural machine translation with smt as posterior regularization. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 241–248, 2019.
  251. 251.Adithya Renduchintala and Adina Williams. Investigating failures of automatic translation in the case of unambiguous gender. CoRR, abs/2104.07838, 2021. URL https://arxiv.org/abs/2104.07838.
  252. 252.Philip Resnik. Mining the Web for Bilingual Text. In ACL, 1999. URL http://www.aclweb.org/anthology/P99-1068.
  253. 253.Philip Resnik and Noah A. Smith. The Web as a Parallel Corpus. Computational Linguistics, 29(3):349–380, 2003. URL http://www.aclweb.org/anthology/J03-3002.
  254. 254.Felix Richter. Infographic: English is the internet’s universal language, February 2022. URL https://www.statista.com/chart/26884/languages-on-the-internet/.
  255. 255.John R. Rickford. Standard and non-standard language attitudes in a creole continuum. In Nessa Wolfson and Joan Manes, editors, Language of Inequality, pages 145–160. De Gruyter Mouton, 2012. doi: doi:10.1515/9783110857320.145. URL https://doi.org/10.1515/9783110857320.145.
  256. 256.Parker Riley, Isaac Caswell, Markus Freitag, and David Grangier. Translationese as a language in “multilingual” nmt. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7737–7746, 2020.
  257. 257.Samantha Robertson, Wesley Hanwen Deng, Timnit Gebru, Margaret Mitchell, Daniel J Liebling, Michal Lahav, Katherine Heller, Mark Díaz, Samy Bengio, and Niloufar Salehi. Three directions for the design of human-centered machine translation. Google Research, 2021.
  258. 258.Björn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki. Measuring the reliability of hate speech annotations: The case of the european refugee crisis. CoRR, abs/1701.08118, 2017. URL http://arxiv.org/abs/1701.08118.
  259. 259.Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, and Fahim Dalvi. AraBench: Benchmarking dialectal Arabic-English machine translation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 5094–5107, Barcelona, Spain (Online), December 2020. International Committee on Computational Linguistics. doi: 10.18653/v1/2020.coling-main.447. URL https://aclanthology.org/2020.coling-main.447.
  260. 260.Mohammad Salameh, Houda Bouamor, and Nizar Habash. Fine-grained Arabic dialect identification. In Coling, pages 1332–1344, Santa Fe, New Mexico, USA, August 2018. Association for Computational Linguistics. URL https://aclanthology.org/C18-1113.
  261. 261.Fahimeh Saleh, Wray Buntine, and Gholamreza Haffari. Collective wisdom: Improving low-resource neural machine translation using adaptive knowledge distillation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 3413–3421, 2020.
  262. 262.Julia Sallabank. Attitudes to endangered languages: Identities and policies. Cambridge University Press, 2013.
  263. 263.Nithya Sambasivan. Seeing like a dataset from the global south. Interactions, 28(4):76–78, 2021.
  264. 264.Nithya Sambasivan and Jess Holbrook. Toward responsible AI for the next billion users. Interactions, 26(1):68–71, 2018.
  265. 265.Nithya Sambasivan, Erin Arnesen, Ben Hutchinson, Tulsee Doshi, and Vinodkumar Prabhakaran. Re-imagining algorithmic fairness in india and beyond. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 315–328, 2021.
  266. 266.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and A Noah Smith. The risk of racial bias in hate speech detection. In ACL, 2019.
  267. 267.Kevin P Scannell. The Crúbadán Project: Corpus building for under-resourced languages. Cahiers du Cental, 5:1, 2007.
  268. 268.Holger Schwenk. Investigations on large-scale lightly-supervised training for statistical machine translation. In Proceedings of the 5th International Workshop on Spoken Language Translation: Papers, 2008.
  269. 269.Holger Schwenk. Filtering and mining parallel data in a joint multilingual space. In ACL, pages 228–234, 2018.
  270. 270.Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. WikiMatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia. In ACL, pages 1351–1361, 2021a.
  271. 271.Holger Schwenk, Guillaume Wenzek, Sergey Edunov, Édouard Grave, Armand Joulin, and Angela Fan. CCMatrix: Mining billions of high-quality parallel sentences on the web. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6490–6500, 2021b.
  272. 272.Thibault Sellam, Dipanjan Das, and Ankur Parikh. Bleurt: Learning robust metrics for text generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7881–7892, 2020.
  273. 273.Rico Sennrich and Biao Zhang. Revisiting low-resource neural machine translation: A case study. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 211–221, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1021. URL https://www.aclweb.org/anthology/P19-1021.
  274. 274.Rico Sennrich, Barry Haddow, and Alexandra Birch. Improving neural machine translation models with monolingual data. Conference of the Association for Computational Linguistics (ACL), 2016a.
  275. 275.Rico Sennrich, Barry Haddow, and Alexandra Birch. Edinburgh neural machine translation systems for wmt 16. In Proceedings of the First Conference on Machine Translation: Volume 2, Shared Task Papers, pages 371–376, 2016b.
  276. 276.Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In Proceedings of International Conference on Learning Representations (ICLR), 2017. URL https://openreview.net/pdf?id=B1ckMDqlg.
  277. 277.Shashi Shekhar, Dilip Kumar Sharma, and MM Sufyan Beg. Language identification framework in code-mixed social media text based on quantum lstm—the word belongs to which language? Modern Physics Letters B, 34(06):2050086, 2020.
  278. 278.Aditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Xu Chen, Sneha Kudugunta, Naveen Arivazhagan, and Yonghui Wu. Leveraging monolingual data with self-supervision for multilingual neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2827–2835, 2020.
  279. 279.Aditya Siddhant, Ankur Bapna, Orhan Firat, Yuan Cao, Mia Xu Chen, Isaac Caswell, and Xavier Garcia. Towards the next 1000 languages in multilingual machine translation: Exploring the synergy between supervised and self-supervised learning. CoRR, abs/2201.03110, 2022. URL https://arxiv.org/abs/2201.03110.
  280. 280.Kathleen Siminyu, Godson Kalipe, Davor Orlic, Jade Abbott, Vukosi Marivate, Sackey Freshia, Prateek Sibal, Bhanu Neupane, David I. Adelani, Amelia Taylor, Jamiil Toure ALI, Kevin Degila, Momboladji Balogoun, Thierno Ibrahima DIOP, Davis David, Chayma Fourati, Hatem Haddad, and Malek Naski. AI4D – african language program. arXiv preprint arXiv:2104.02516, 2021.
  281. 281.Nitish Singh, Kevin Lehnert, and Kathleen Bostick. Global social media usage: Insights into reaching consumers worldwide. Thunderbird International Business Review, 54(5): 683–700, 2012.
  282. 282.Raivis Skadiņš, Mārcis Pinnis, Andrejs Vasiļjevs, Inguna Skadiņa, and Tomas Hudik. Application of machine translation in localization into low-resourced languages. In Proceedings of the 17th Annual conference of the European Association for Machine Translation, pages 209–216, Dubrovnik, Croatia, June 2014a. European Association for Machine Translation. URL https://aclanthology.org/2014.eamt-1.43.
  283. 283.Raivis Skadiņš, Jörg Tiedemann, Roberts Rozis, and Daiga Deksne. Billions of parallel words for free: Building and using the EU bookshop corpus. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pages 1850–1855, Reykjavik, Iceland, May 2014b. European Language Resources Association (ELRA). URL http://www.lrec-conf.org/proceedings/lrec2014/pdf/846_Paper.pdf.
  284. 284.Lucia Specia, Frédéric Blain, Marina Fomicheva, Chrysoula Zerva, Zhenhao Li, Vishrav Chaudhary, and André F. T. Martins. Findings of the WMT 2021 shared task on quality estimation. In Proceedings of the Sixth Conference on Machine Translation, pages 684–725, Online, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.wmt-1.71.
  285. 285.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014.
  286. 286.Haipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, and Tiejun Zhao. Knowledge distillation for multilingual unsupervised neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3525–3535, 2020.
  287. 287.Cass Robert Sunstein and Richard Thaler. Libertarian paternalism. American Economic Review, 2003.
  288. 288.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. CoRR, abs/1512.00567, 2015. URL http://arxiv.org/abs/1512.00567.
  289. 289.Xu Tan, Yi Ren, Di He, Tao Qin, Zhou Zhao, and Tie-Yan Liu. Multilingual neural machine translation with knowledge distillation. In International Conference on Learning Representations, 2018.
  290. 290.Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. Multilingual translation with extensible multilingual pretraining and finetuning. CoRR, abs/2008.00401, 2020. URL https://arxiv.org/abs/2008.00401.
  291. 291.Martin Thoma. The wili benchmark dataset for written language identification. CoRR, abs/1801.07779, 2018. URL http://arxiv.org/abs/1801.07779.
  292. 292.J. Tiedemann. Parallel data, tools and interfaces in OPUS. In LREC, 2012.
  293. 293.Chau Tran, Shruti Bhosale, James Cross, Philipp Koehn, Sergey Edunov, and Angela Fan. Facebook AI’s WMT21 news translation task submission. In Proceedings of the Sixth Conference on Machine Translation, pages 205–215, Online, November 2021. Association for Computational Linguistics. URL https://aclanthology.org/2021.wmt-1.19.
  294. 294.Emiliano Treré. The dark side of digital politics: Understanding the algorithmic manufacturing of consent and the hindering of online dissidence. Opening Governance, 2016.
  295. 295.Masao Utiyama and Hitoshi Isahara. Reliable Measures for Aligning Japanese-English News Articles and Sentences. In ACL, 2003. URL http://www.aclweb.org/anthology/P03-1010.
  296. 296.Jeroen Van Der Hoven and Noemi Manders-Huits. Value-sensitive design. In The Ethics of Information Technologies, pages 329–332. Routledge, 2020.
  297. 297.Jack Vance. The Eyes of the overworld, volume 3. Macmillan Reference USA, 1977.
  298. 298.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  299. 299.Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts. Challenges and frontiers in abusive content detection. In Proceedings of the third workshop on abusive language online, pages 80–93. Association for Computational Linguistics, 2019.
  300. 300.Vered Volansky, Noam Ordan, and Shuly Wintner. On the features of translationese. Digital Scholarship in the Humanities, 30(1):98–118, 2015.
  301. 301.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations, 2018.
  302. 302.Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, and Furu Wei. Deepnet: Scaling transformers to 1,000 layers, 2022. URL https://arxiv.org/abs/2203.00555.
  303. 303.Yiren Wang, ChengXiang Zhai, and Hany Hassan. Multi-task learning for multilingual neural machine translation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1022–1034, Online, November 2020a. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.75. URL https://www.aclweb.org/anthology/2020.emnlp-main.75.
  304. 304.Zihan Wang, Karthikeyan K, Stephen Mayhew, and Dan Roth. Extending multilingual bert to low-resource languages, 2020b. URL https://arxiv.org/abs/2004.13640.
  305. 305.Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao. Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models. In International Conference on Learning Representations, 2020c.
  306. 306.Steven Weber. The success of open source. Harvard University Press, 2004.
  307. 307.Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. Challenges in detoxifying language models. CoRR, abs/2109.07445, 2021. URL https://arxiv.org/abs/2109.07445.
  308. 308.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Édouard Grave. CCNet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of The 12th Language Resources and Evaluation Conference, pages 4003–4012, 2020.
  309. 309.Dominic Widdows and Chris Brew. Language identification with a reciprocal rank classifier. CoRR, abs/2109.09862, 2021. URL https://arxiv.org/abs/2109.09862.
  310. 310.Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al. Sustainable ai: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4:795–813, 2022.
  311. 311.Shijie Wu, Ryan Cotterell, and Mans Hulden. Applying the transformer to character-level transduction. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1901–1907, Online, April 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.163. URL https://aclanthology.org/2021.eacl-main.163.
  312. 312.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. Google’s neural machine translation system: Bridging the gap between human and machine translation. CoRR, abs/1609.08144, 2016. URL http://arxiv.org/abs/1609.08144.
  313. 313.Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pages 10524–10533. PMLR, 2020.
  314. 314.Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein. Detoxifying language models risks marginalizing minority voices. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2390–2397, Online, June 2021a. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.190. URL https://aclanthology.org/2021.naacl-main.190.
  315. 315.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. Recipes for safety in open-domain chatbots. CoRR, abs/2010.07079, 2020. URL https://arxiv.org/abs/2010.07079.
  316. 316.Jing Xu, Arthur Szlam, and Jason Weston. Beyond goldfish memory: Long-term open-domain conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5180–5197, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.356. URL https://aclanthology.org/2022.acl-long.356.
  317. 317.Weijia Xu, Shuming Ma, Dongdong Zhang, and Marine Carpuat. How does distilled data complexity impact the quality and confidence of non-autoregressive machine translation? In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4392–4400, 2021b.
  318. 318.Yinfei Yang, Gustavo Hernandez Abrego, Steve Yuan, Mandy Guo, Qinlan Shen, Daniel Cer, Yun-hsuan Sung, Brian Strope, and Ray Kurzweil. Improving multilingual sentence embedding using bi-directional dual encoder with additive margin softmax. In IJCAI, pages 5370–5378, 2019.
  319. 319.Zhen Yang, Wei Chen, Feng Wang, and Bo Xu. Unsupervised neural machine translation with weight sharing. CoRR, abs/1804.09057, 2018. URL http://arxiv.org/abs/1804.09057.
  320. 320.Marcos Zampieri, Binyam Gebrekidan Gebre, Hernani Costa, and Josef Van Genabith. Comparing approaches to the identification of similar languages. In Proceedings of the Joint Workshop on Language Technology for Closely Related Languages, Varieties and Dialects, pages 66–72, 2015a.
  321. 321.Marcos Zampieri, Liling Tan, Nikola Ljubešić, Jörg Tiedemann, and Preslav Nakov. Overview of the DSL shared task 2015. In Proceedings of the Joint Workshop on Language Technology for Closely Related Languages, Varieties and Dialects, pages 1–9, 2015b.
  322. 322.Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. SemEval-2019 task 6: Identifying and categorizing offensive language in social media (OffensEval). In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 75–86, Minneapolis, Minnesota, USA, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/S19-2010. URL https://aclanthology.org/S19-2010.
  323. 323.Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. Improving massively multilingual neural machine translation and zero-shot translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1628–1639, 2020.
  324. 324.Biao Zhang, Ankur Bapna, Rico Sennrich, and Orhan Firat. Share or not? learning to schedule language-specific capacity for multilingual translation. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=Wj4ODo0uyCF.
  325. 325.Dakun Zhang, Josep M Crego, and Jean Senellart. Analyzing knowledge distillation in neural machine translation. In Proceedings of the 15th International Conference on Spoken Language Translation, pages 23–30, 2018.
  326. 326.Mike Zhang and Antonio Toral. The effect of translationese in machine translation test sets. arXiv preprint arXiv:1906.08069, 2019.
  327. 327.Chunting Zhou, Graham Neubig, and Jiatao Gu. Understanding knowledge distillation in non-autoregressive machine translation. CoRR, abs/1911.02727, 2020. URL http://arxiv.org/abs/1911.02727.
  328. 328.Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. Transfer learning for low-resource neural machine translation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1568–1575, Austin, Texas, November 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1163. URL https://aclanthology.org/D16-1163.
  329. 329.Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. Designing effective sparse expert models. arXiv preprint arXiv:2202.08906, 2022.
  330. 330.Shoshana Zuboff. The age of surveillance capitalism: The fight for a human future at the new frontier of power: Barack Obama’s books of 2019. Profile books, 2019.
  331. 331.Ethan Zuckerman. The polyglot internet, October 2008. URL https://ethanzuckerman.com/the-polyglot-internet/.
  332. 332.Alp Öktem, Muhannad Albayk Jaam, Eric DeLuca, and Grace Tang. Gamayun – language technology for humanitarian response. In 2020 IEEE Global Humanitarian Technology Conference (GHTC), pages 1–4, 2020. doi: 10.1109/GHTC46280.2020.9342939.

Citation

MLA
Team, N., et al. “No Language Left Behind: Scaling Human-Centered Machine Translation”. arXiv, 2022, http://arxiv.org/abs/2207.04672v3.
APA
Team, N., Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., Sun, A., Wang, S., Wenzek, G., Youngblood, A., Akula, B., Barrault, L., Gonzalez, G. M., Hansanti, P., … Wang, J. (2022). No Language Left Behind: Scaling Human-Centered Machine Translation. arXiv. http://arxiv.org/abs/2207.04672v3
Chicago
Team, N., M. R. Costa-jussà, J. Cross, et al. 2022. “No Language Left Behind: Scaling Human-Centered Machine Translation”. arXiv. http://arxiv.org/abs/2207.04672v3.
Harvard
Team, N. et al. (2022) “No Language Left Behind: Scaling Human-Centered Machine Translation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2207.04672v3.
Vancouver
1. Team N, Costa-jussà MR, Cross J, et al (2022) No Language Left Behind: Scaling Human-Centered Machine Translation. arXiv

BibTeX

@article{team2022language,
  title = {No Language Left Behind: Scaling Human-Centered Machine Translation},
  author = {Team, NLLB and Costa-jussà, Marta R. and Cross, James and Çelebi, Onur and Elbayad, Maha and Heafield, Kenneth and Heffernan, Kevin and Kalbassi, Elahe and Lam, Janice and Licht, Daniel and Maillard, Jean and Sun, Anna and Wang, Skyler and Wenzek, Guillaume and Youngblood, Al and Akula, Bapi and Barrault, Loic and Gonzalez, Gabriel Mejia and Hansanti, Prangthip and Hoffman, John and Jarrett, Semarley and Sadagopan, Kaushik Ram and Rowe, Dirk and Spruit, Shannon and Tran, Chau and Andrews, Pierre and Ayan, Necip Fazil and Bhosale, Shruti and Edunov, Sergey and Fan, Angela and Gao, Cynthia and Goswami, Vedanuj and Guzmán, Francisco and Koehn, Philipp and Mourachko, Alexandre and Ropers, Christophe and Saleem, Safiyyah and Schwenk, Holger and Wang, Jeff},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2207.04672v3},
  eprint = {2207.04672}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by-sa/4.0/