The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

Lucas BandarkarDavis LiangBenjamin MullerMikel ArtetxeSatya Narayan ShuklaDonald HusaNaman GoyalAbhinandan KrishnanLuke ZettlemoyerMadian Khabsa

article2024ACL385 citations

Introduces a fully parallel reading comprehension benchmark across 122 language variants to evaluate and directly compare language models across high-, medium-, and low-resource languages.

Listen

As artificial intelligence applications expand globally, natural language processing systems face growing demands to support diverse global populations. However, the lack of high-quality, parallel evaluation datasets across a broad range of languages severely limits the assessment and development of technologies for non-English speakers, particularly in medium- and low-resource languages. Most existing language understanding benchmarks cover fewer than 30 languages, leaving critical blind spots regarding whether state-of-the-art models truly generalize across global linguistic boundaries.

To address this critical evaluation gap, the article introduces BELEBELE, a parallel multiple-choice reading comprehension benchmark designed to evaluate language models across 122 language variants. The primary objective is to measure, compare, and establish baseline performance for both multilingual masked language models and English-centric large language models across diverse language families and resource levels.

The benchmark consists of 900 curated four-option multiple-choice questions derived from 488 passages sourced from the FLORES-200 machine translation corpus, spanning 27 language families and 29 scripts. Professional human translators created the dataset without machine translation, ensuring full parallelism across all 122 language variants, including romanized variants of five Indo-Aryan languages. The authors designed questions to measure text comprehension while resisting simple pattern-matching shortcuts, confirmed via iterative quality checks and statistical featurization. The evaluation assessed multiple masked language models (such as XLM-V and INFOXLM) and prominent large language models (including LLAMA variants, FALCON, and GPT-3.5-Turbo) across few-shot, zero-shot, and fine-tuning configurations.

The analysis revealed several critical findings regarding language model capabilities. First, masked language models trained on balanced multilingual data with extensive vocabularies vastly outperform English-centric large language models on language coverage. For instance, XLM-V achieved above 50% accuracy on over 76% of evaluated languages when fine-tuned, whereas GPT-3.5-Turbo exceeded that threshold on only 43.4% of languages and LLAMA 2 (70B) on 38.5%. Second, English-centric models excel on top-tier languages—such as LLAMA 2 (70B) reaching 90.9% accuracy on English—but drop off sharply on medium- and low-resource languages. Third, cascading translation provides a substantial operational benefit for large language models: translating non-English inputs into English before inference (Translate-Test) raised LLAMA-2-CHAT (70B) average zero-shot accuracy from 44.0% to 57.1% across 91 languages, improving performance in 68 languages. Finally, human evaluation achieved 97.6% accuracy on English questions, substantially higher than all evaluated models, demonstrating that the benchmark presents a demanding, meaningful challenge.

These findings have major strategic implications for deploying language technologies across international markets. Relying solely on raw, English-centric foundation models introduces severe operational and performance risks in non-English contexts, risking total failure in low-resource environments. Vocabulary size and pretraining data distribution represent foundational bottlenecks; models with larger, tailored tokenizers such as XLM-V retain linguistic capabilities much deeper into the long tail of global languages. In addition, the success of translating inputs into English demonstrates that pipeline architectures combining specialized machine translation with large language models offer a cost-effective alternative to training massive native models for each language.

Organizations and developers should adopt specific, evidence-backed practices based on these results. When serving global user bases, teams should deploy architectures that leverage machine translation pipelines or specialized multilingual models rather than relying on direct zero-shot prompting of English-centric systems. Model builders should also allocate larger vocabulary budgets and include balanced multilingual corpora during pretraining. Looking ahead, future research should develop complementary benchmarks that measure nuanced cultural factors—such as local formality and societal values—which pure parallel translations do not capture.

The conclusions should be interpreted in light of certain limitations. The benchmark is derived from English source text and parallel translations, which introduces subtle translation artifacts and excludes localized cultural context. Furthermore, lack of transparency regarding the exact training data compositions for proprietary models like GPT-3.5-Turbo limits full experimental reproducibility. Nevertheless, the rigorous multi-stage human validation and statistical quality checks support high confidence in the benchmark’s comparative evaluations across all 122 language variants.

arXiv: 2308.16884facebookresearch/belebele
Cover for The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

Abstract

We present Belebele, a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. Significantly expanding the language coverage of natural language understanding (NLU) benchmarks, this dataset enables the evaluation of text models in high-, medium-, and low-resource languages. Each question is based on a short passage from the FLORES-200 dataset and has four multiple-choice answers. The questions were carefully curated to discriminate between models with different levels of general language comprehension. The English dataset on its own proves difficult enough to challenge state-of-the-art language models. Being fully parallel, this dataset enables direct comparison of model performance across all languages. We use this dataset to evaluate the capabilities of multilingual masked language models (MLMs) and large language models (LLMs). We present extensive results and findings, notably that despite significant cross-lingual transfer in English-centric LLMs, much smaller MLMs pretrained on balanced multilingual data still understand far more languages. Overall, Belebele opens up new avenues for evaluating and analyzing the multilingual capabilities of NLP systems.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Cross-Lingual Evaluation Benchmarks
  • 2.2 Non-English Machine Reading Comprehension
  • 2.3 Multiple Choice QA
  • 2.4 FLORES-200
  • 3 The BELEBELE Dataset
  • 3.1 Creation of Multiple Choice Questions & Answers
  • 3.2 Quality Assurance
  • 3.3 Translating the Corpus
  • 3.4 English Training Data
  • 3.5 The BELEBELE Dataset in Summary
  • 4 Experiments
  • 4.1 Evaluated Models
  • 4.2 Evaluation Settings
  • 5 Results
  • 5.1 How difficult is BELEBELE?
  • 5.2 Multilingual Generalization of MLMs and LLMs on BELEBELE
  • Scaling effect on Multilingual Generalization
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • References
  • A Appendix
  • A.1 Languages and Variants
  • A.2 Training Set
  • A.3 Licensing
  • A.4 Experiment Details
  • A.4.1 Model fine-tuning
  • A.4.2 In-Context Learning Prompt
  • A.4.3 Zero-Shot Instructions
  • A.5 Annotation Guidelines
  • A.5.1 MCQA Annotation Guidelines
  • A.5.2 Translation Specifications
  • A.6 MCQA Lexical Featurization
  • A.7 Detailed Results Tables
  • A.7.1 Cross-Lingual MLMs
  • A.7.2 LLMs
  • A.7.3 Languages in Multiple Scripts
  • A.7.4 Translate-Test

Knowls

  1. Knowl 1 — BELEBELE Benchmark Dataset Structure and Statistics

    definition

    BELEBELE is a fully parallel, 4-choice multiple-choice reading comprehension (MRC) dataset spanning 122 language variants (115 distinct languages when ignoring script differences), 29 distinct scripts, and 27 language families.

    The dataset contains 900 unique questions linked to 488 distinct short passages sourced from the non-test split of the FLORES-200 machine translation benchmark (derived from Wikinews, Wikijunior, and WikiVoyage). Each passage is accompanied by 1 to 2 questions, and each question has exactly one correct answer and three distractor choices. Fully expanded across all 122 language variants, the corpus comprises 109,800 examples.

    Key length statistics measured on the English split are:

    • Average passage length: 79.1±26.279.1 \pm 26.2 words (4.1±1.44.1 \pm 1.4 sentences)
    • Average question length: 12.9±4.012.9 \pm 4.0 words
    • Average candidate answer length: 4.2±2.94.2 \pm 2.9 words

    The benchmark includes parallel romanized (Latin script) variants alongside native-script variants for five Indo-Aryan languages (Hindi, Bengali, Urdu, Nepali, and Sinhala) as well as Modern Standard Arabic, and both Simplified and Traditional scripts for Chinese.

  2. Knowl 2 — Question Construction and Heuristic Bias Filtering Protocol

    model/method

    BELEBELE questions and candidate answers were constructed in English by professional Language Service Provider (LSP) annotators over five iterative feedback cycles, followed by human translation into 121 target language variants.

    To ensure questions test text comprehension rather than superficial pattern-matching or reasoning shortcuts, candidate items were screened through manual validation and programmatic lexical featurization based on 1-, 2-, 3-, and 4-gram overlap:

    1. Lexical overlap between passage and correct answer vs. wrong answers (preventing extraction-only heuristics).
    2. Lexical overlap between question and correct answer vs. wrong answers.
    3. Sentence-level colocation: whether the correct answer appears in the exact sentence overlapping with the question.
    4. Lexical similarity between the correct answer and distractor choices.
    5. Answer length statistics (character and word counts).

    Statistical tt-tests confirmed that the feature distributions for correct vs. incorrect answers in the final dataset are indistinguishable (p=0.81p = 0.81, compared to p<0.01p < 0.01 for MCTest). A naive bag-of-words logistic regression baseline achieved an accuracy of only 0.280.28 (near the 0.250.25 random guessing baseline, compared to 0.440.44 on MCTest). Approximately 20% of candidate items failing heuristic quality thresholds were filtered out during the final curation iteration.

  3. Knowl 3 — English Multi-Source Training Set for Multiple-Choice QA Fine-Tuning

    model/method

    Because BELEBELE is designed exclusively as an evaluation benchmark without an official training split, an external English multiple-choice training and development set was constructed by pooling and standardizing instances from six existing MRC benchmarks: RACE, SCIQ, MULTIRC, MCTEST, MCSCRIPT2.0, and RECLOR.

    Passages and questions were re-formatted into a uniform 4-choice single-answer format (e.g., sampling distractors from other questions for two-choice datasets like MCScript2.0, and filtering out multi-answer or fill-in-the-blank instances). Subgroups of data were filtered and stratified based on passage length, question length, and topic, optimizing validation performance on a RoBERTa-base model evaluated on the BELEBELE English set.

    The final standardized training set consists of 67.5k training examples (over half sourced from RACE) and 3.7k validation examples. For multilingual "Translate-Train-All" MLM fine-tuning, this dataset was machine-translated across target languages and subsampled to a cap of 650k examples.

  4. Knowl 4 — Comparative Performance of Multilingual MLMs and Large Language Models on BELEBELE

    data/table

    Language models evaluated on BELEBELE exhibit a stark trade-off: English-centric LLMs (such as LLaMA 1, LLaMA 2, and Falcon 40B) attain higher peak accuracy on English and major high-resource languages, but multilingual Masked Language Models (MLMs) pretrained on balanced corpora (XLM-R, XLM-V, and InfoXLM) generalize significantly better across medium- and low-resource languages.

    Model Size AVG % ≥\ge 50 % ≥\ge 70 eng_Latn non-Eng AVG
    5-Shot In-Context Learning (English examples)
    LLAMA 1 7B 27.7 0.0% 0.0% 37.3 27.6
    LLAMA 1 13B 30.4 0.8% 0.0% 53.3 30.2
    LLAMA 1 30B 36.2 18.0% 0.8% 73.1 35.9
    LLAMA 1 70B 40.9 25.4% 12.3% 82.5 40.5
    LLAMA 2 base 70B 48.0 38.5% 26.2% 90.9 47.7
    FALCON 40B 37.3 16.4% 1.6% 77.2 36.9
    Zero-Shot for Instructed Models (English instructions)
    LLAMA-2-CHAT 7B 34.4 4.1% 0.0% 58.6 34.1
    LLAMA-2-CHAT 70B 41.5 27.0% 2.5% 78.8 41.2
    GPT3.5-TURBO unk 51.1 44.2% 29.2% 87.7 50.7
    Full Finetuning in English (Zero-Shot Transfer)
    XLM-R large (550M) 54.0 64.8% 15.6% 76.2 53.8
    XLM-V large (1.2B) 55.6 69.7% 21.2% 76.2 54.9
    INFOXLM large (550M) 56.2 67.2% 28.7% 79.3 56.0
    Translate-Train-All
    XLM-R large (550M) 58.9 69.7% 36.1% 78.7 58.8
    XLM-V large (1.2B) 60.2 76.2% 32.8% 77.8 60.1
    INFOXLM large (550M) 60.0 70.5% 36.9% 81.2 59.8

    Metrics:

    • AVG\text{AVG}: Macro-average accuracy across all 122 language variants (random baseline is 25.0%).
    • %≥50\% \ge 50 / %≥70\% \ge 70: Proportion of the 122 language variants where the model scores ≥50%\ge 50\% or ≥70%\ge 70\% accuracy.
    • eng_Latn\text{eng\_Latn}: Accuracy on the English split.
  5. Knowl 5 — Translate-Test Cascading vs. Direct In-Language Prompting in LLMs

    empirical result

    Cascading machine translation into English before inference (Translate-Test) substantially outperforms direct in-language prompting for English-centric LLMs on medium- and low-resource languages.

    When evaluating LLAMA-2-CHAT (70B) in zero-shot across 91 non-English languages:

    • Translate-Test (translating target passages, questions, and options into English) achieved an average accuracy of 57.1%57.1\% and scored ≥50%\ge 50\% accuracy on 78.0%78.0\% of languages (71 of 91 languages).
    • Direct In-Language Evaluation achieved an average accuracy of 44.1%44.1\% and scored ≥50%\ge 50\% accuracy on only 35.2%35.2\% of languages (32 of 91 languages).
    • Translate-Test outperformed direct in-language evaluation on 68 of the 91 languages, with improvements exceeding 20 percentage points on nearly all low-resource languages. Only two high-resource languages performed non-trivially better natively in-language: German (69.4%69.4\% in-language vs. 65.7%65.7\% Translate-Test) and Italian (68.6%68.6\% in-language vs. 66.1%66.1\% Translate-Test).

    Additionally, machine-translating the task prompt instructions into the target language degraded zero-shot performance compared to providing English instructions (38.7%38.7\% average vs. 44.9%44.9\% on 89 languages), with instructions failing entirely (below random chance) for roughly 25 low-scoring languages.

  6. Knowl 6 — Subword Vocabulary Size Impact on Low-Resource Multilingual NLU

    empirical result

    Subword vocabulary capacity is a key structural determinant of model generalization on low-resource language comprehension.

    Among masked language models trained on the CC-100 corpus with identical transformer architectures:

    • XLM-V (vocabulary size of 902K tokens) outperforms XLM-R and INFOXLM (each with 250K token vocabularies) on the tail of lower-resource languages. Under the Translate-Train-All regime, XLM-V achieves ≥50%\ge 50\% accuracy on 76.2%76.2\% of the 122 languages, compared to 70.5%70.5\% for INFOXLM and 69.7%69.7\% for XLM-R.
    • While INFOXLM achieves higher peak accuracy on top high- and medium-resource languages (37.2%37.2\% of languages ≥70%\ge 70\% vs. 33.1%33.1\% for XLM-V), XLM-V's language-dedicated token allocation prevents out-of-vocabulary degradation in low-resource settings.
    • In LLMs, models with small vocabularies (e.g., LLAMA 1 and LLAMA 2 with 32K vocabularies) suffer steep performance degradation on non-English languages, whereas FALCON 40B (65K vocabulary) achieves multilingual performance on par with LLAMA 1 30B despite Falcon having a lower percentage of non-English pretraining data.
  7. Knowl 7 — Parameter Scale and Cross-Lingual Transfer Across Language Families

    empirical result

    Model parameter scaling is critical for eliciting cross-lingual reading comprehension capabilities in English-centric autoregressive models, especially for distant language families.

    Evaluating LLAMA 1 across parameter checkpoints (7B, 13B, 30B, and 65B) in 5-shot in-context learning shows:

    • The 7B checkpoint performs barely above random guessing (25.0%25.0\%) across non-English language families and scores only 37.3%37.3\% on English.
    • Cross-lingual transfer scales monotonically with model size across all language families, including Romance (6 languages), Germanic (8 languages), Austronesian (10 languages), and Dravidian (4 languages).
    • For language families unrepresented in LLAMA 1's documented pretraining corpus (such as Japonic and Hellenic), performance remains near random chance at 7B and 13B, and only exhibits non-trivial comprehension above chance at the 30B and 65B scales.
  8. Knowl 8 — Native Script vs. Romanized Script Comprehension Disparity

    empirical result

    Evaluating multilingual models on paired native-script vs. Romanized (Latin script) transliterations of identical passages reveals that models generally comprehend native scripts more effectively than Romanized transliterations.

    Across Modern Standard Arabic and five Indo-Aryan languages (Hindi, Bengali, Nepali, Sinhala, Urdu):

    • Modern Standard Arabic: GPT-3.5-Turbo scores 69.3%69.3\% in Arabic script (arb_Arab) vs. 31.1%31.1\% in Latin script (arb_Latn); InfoXLM (fine-tuned in English) scores 71.0%71.0\% in Arabic script vs. 32.2%32.2\% in Latin script.
    • Bengali: InfoXLM scores 63.4%63.4\% in Bengali script (ben_Beng) vs. 36.9%36.9\% in Latin script (ben_Latn).
    • Hindi: LLAMA 2 (70B, 5-shot) scores 52.6%52.6\% in Devanagari script (hin_Deva) vs. 49.0%49.0\% in Latin script (hin_Latn); InfoXLM scores 60.2%60.2\% in Devanagari vs. 49.7%49.7\% in Latin script.
    • Exception: FALCON 40B exhibits the reverse trend for all five Indo-Aryan languages, performing higher on Romanized text than on native scripts (e.g., Hindi: 40.0%40.0\% Latin vs. 27.1%27.1\% Devanagari; Urdu: 34.2%34.2\% Latin vs. 31.7%31.7\% Arabic script; Sinhala: 32.6%32.6\% Latin vs. 27.7%27.7\% Sinhala script).
  9. Knowl 9 — Human Performance Ceiling and Model Difficulty on BELEBELE

    empirical result

    BELEBELE English reading comprehension questions present a substantial gap between human understanding and state-of-the-art NLP models:

    • In a blind evaluation of randomly sampled English items, human annotators achieved a mean accuracy of 97.6%97.6\% (with a 95% confidence interval of [93.1%,99.5%][93.1\%, 99.5\%]).
    • Fully fine-tuned RoBERTa-base achieved an English accuracy of 71.7%71.7\%.
    • Zero-shot LLAMA-2-CHAT (70B) reached 78.8%78.8\%, GPT-3.5-Turbo reached 87.7%87.7\%, and 5-shot LLAMA 2 (70B) reached 90.9%90.9\%.
    • Model accuracy on BELEBELE exhibits a strong linear correlation with performance on the XNLI benchmark (r=0.85r = 0.85 across 15 overlapping languages for XLM-R, XLM-V, and InfoXLM), while scoring systematically ∼10\sim 10 accuracy points lower on BELEBELE than on XNLI under Translate-Train.
  10. Knowl 10 — Methodological Limitations of Parallel Translation-Based NLU Evaluation

    limitation

    The construction methodology and evaluation design of BELEBELE entail specific limitations:

    1. Western-Centrism and Translationese: By translating English source passages and questions into 121 target language variants to preserve cross-lingual comparability, the benchmark inevitably contains translation artifacts ("translationese") and cannot capture language-specific or culturally situated phenomena such as culture-specific reasoning, localized formality levels, and culture-specific values.
    2. Inherited Base Translation Errors: Because passages were drawn directly from FLORES-200, occasional dialectal or translation inaccuracies in low-resource FLORES passages constrained the phrasing and alignment of target questions and answers.
    3. Pretraining Data Opacity: Lack of transparency and detailed language composition reporting in proprietary pretraining corpora (e.g., GPT-3.5-Turbo) prevents exact attribution of zero-shot capabilities to true zero-shot cross-lingual transfer versus pretraining corpus leakage.

Coverage note — None was omitted; all key contributions including dataset statistics, annotation protocols, fine-tuning setups, empirical comparisons across MLMs and LLMs, translation cascading experiments, vocabulary/scaling/script analyses, human evaluations, and limitations are fully covered.

References

  1. 1.Manish Agarwal and Prashanth Mannem. 2011. Automatic gap-fill question generation from text books. In Proceedings of the Sixth Workshop on Innovative Use of NLP for Building Educational Applications, pages 56–64, Portland, Oregon. Association for Computational Linguistics.
  2. 2.Kaveri Anuranjana, Vijjini Anvesh Rao, and Radhika Mamidi. 2019. Hindirc: A dataset for reading comprehension in hindi. In 20th International Conference on Computational Linguistics and Intelligent Text.
  3. 3.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637, Online. Association for Computational Linguistics.
  4. 4.Emily M. Bender. 2009. Linguistically naïve != language independent: Why NLP needs linguistic typology. In Proceedings of the EACL 2009 Workshop on the Interaction between Linguistics and Computational Linguistics: Virtuous, Vicious or Vacuous?, pages 26–32, Athens, Greece. Association for Computational Linguistics.
  5. 5.Emily M. Bender. 2011. On achieving and evaluating language-independence in nlp. Linguistic Issues in Language Technology, 6.
  6. 6.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146.
  7. 7.Samuel R. Bowman, Jennimaria Palomaki, Livio Baldini Soares, and Emily Pitler. 2020. New protocols and negative results for textual entailment data collection. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8203–8214, Online. Association for Computational Linguistics.
  8. 8.Jordan Boyd-Graber and Benjamin Börschinger. 2020. What question answering can learn from trivia nerds. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7422–7435, Online. Association for Computational Linguistics.
  9. 9.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  10. 10.Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou. 2021. InfoXLM: An information-theoretic framework for cross-lingual language model pre-training. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3576–3588, Online. Association for Computational Linguistics.
  11. 11.Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8:454–470.
  12. 12.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020a. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguistics.
  13. 13.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485, Brussels, Belgium. Association for Computational Linguistics.
  14. 14.Alexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov. 2020b. Emerging cross-lingual structure in pretrained language models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6022–6034, Online. Association for Computational Linguistics.
  15. 15.Danilo Croce, Alexandra Zelenanska, and Roberto Basili. 2018. Neural learning for question answering in italian. In International Conference of the Italian Association for Artificial Intelligence.
  16. 16.Martin d’Hoffschmidt, Wacim Belblidia, Quentin Heinrich, Tom Brendlé, and Maxime Vidal. 2020. FQuAD: French question answering dataset. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1193–1208, Online. Association for Computational Linguistics.
  17. 17.Pavel Efimov, Andrey Chertok, Leonid Boytsov, and Pavel Braslavski. 2020. SberQuAD – russian reading comprehension dataset: Description and analysis. In Lecture Notes in Computer Science, pages 3–15. Springer International Publishing.
  18. 18.Asım Ersoy, Gerson Vizcarra, Tahsin Mayeesha, and Benjamin Muller. 2023. In what languages are generative language models the most formal? analyzing formality distribution across languages. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 2650–2666, Singapore. Association for Computational Linguistics.
  19. 19.Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. 2023. MASSIVE: A 1M-example multilingual natural language understanding dataset with 51 typologically-diverse languages. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4277–4302, Toronto, Canada. Association for Computational Linguistics.
  20. 20.Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022. The Flores-101 evaluation benchmark for low-resource and multilingual machine translation. Transactions of the Association for Computational Linguistics, 10:522–538.
  21. 21.Deepak Gupta, Surabhi Kumari, Asif Ekbal, and Pushpak Bhattacharyya. 2018. MMQA: A multi-domain multi-lingual question-answering framework for English and Hindi. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).
  22. 22.Momchil Hardalov, Ivan Koychev, and Preslav Nakov. 2019. Beyond English-only reading comprehension: Experiments in zero-shot multilingual transfer for Bulgarian. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2019), pages 447–459, Varna, Bulgaria. INCOMA Ltd.
  23. 23.Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. 2020. EXAMS: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5427–5444, Online. Association for Computational Linguistics.
  24. 24.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. XL-sum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703, Online. Association for Computational Linguistics.
  25. 25.Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022. Challenges and strategies in cross-cultural NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6997–7013, Dublin, Ireland. Association for Computational Linguistics.
  26. 26.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  27. 27.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
  28. 28.Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018. Looking beyond the surface: A challenge set for reading comprehension over multiple sentences. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 252–262, New Orleans, Louisiana. Association for Computational Linguistics.
  29. 29.Grgur Kovač, Masataka Sawayama, Rémy Portelas, Cédric Colas, Peter Ford Dominey, and Pierre-Yves Oudeyer. 2023. Large language models as superpositions of cultural perspectives. arXiv preprint arXiv:2307.07870.
  30. 30.Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, and Mofetoluwa Adeyemi. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, 10:50–72.
  31. 31.Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium. Association for Computational Linguistics.
  32. 32.Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020. WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4034–4048, Online. Association for Computational Linguistics.
  33. 33.Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. RACE: Large-scale ReAding comprehension dataset from examinations. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785–794, Copenhagen, Denmark. Association for Computational Linguistics.
  34. 34.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. MLQA: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7315–7330, Online. Association for Computational Linguistics.
  35. 35.Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023. XLM-V: Overcoming the vocabulary bottleneck in multilingual masked language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13142–13152, Singapore. Association for Computational Linguistics.
  36. 36.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  37. 37.Yash Madhani, Sushane Parthan, Priyanka Bedekar, Gokul Nc, Ruchi Khapra, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Khapra. 2023. Aksharantar: Open Indic-language transliteration datasets and models for the next billion users. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 40–57, Singapore. Association for Computational Linguistics.
  38. 38.Chaitanya Malaviya, Sudeep Bhatia, and Mark Yatskar. 2022. Cascading biases: Investigating the effect of heuristic annotation strategies on data and models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 6525–6540, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  39. 39.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, Brussels, Belgium. Association for Computational Linguistics.
  40. 40.Timo Möller, Julian Risch, and Malte Pietsch. 2021. GermanQuAD and GermanDPR: Improving non-English question answering and passage retrieval. In Proceedings of the 3rd Workshop on Machine Reading for Question Answering, pages 42–50, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  41. 41.Hussein Mozannar, Elie Maamary, Karl El Hajal, and Hazem Hajj. 2019. Neural Arabic question answering. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pages 108–118, Florence, Italy. Association for Computational Linguistics.
  42. 42.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023. Crosslingual generalization through multitask finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15991–16111, Toronto, Canada. Association for Computational Linguistics.
  43. 43.Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021a. When being unseen from mBERT is just the beginning: Handling new languages with multilingual language models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 448–462, Online. Association for Computational Linguistics.
  44. 44.Benjamin Muller, Yanai Elazar, Benoît Sagot, and Djamé Seddah. 2021b. First align, then predict: Understanding the cross-lingual ability of multilingual BERT. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2214–2231, Online. Association for Computational Linguistics.
  45. 45.Benjamin Muller, Deepanshu Gupta, Jean-Philippe Fauconnier, Siddharth Patwardhan, David Vandyke, and Sachin Agarwal. 2023. Languages you know influence those you learn: Impact of language characteristics on multi-lingual text-to-text transfer. In Proceedings of The 1st Transfer Learning for Natural Language Processing Workshop, volume 203 of Proceedings of Machine Learning Research, pages 88–102. PMLR.
  46. 46.Benjamin Muller, Benoît Sagot, and Djamé Seddah. 2020. Can multilingual language models transfer to an unseen dialect? a case study on north african arabizi. ArXiv, abs/2005.00318.
  47. 47.Nikita Nangia and Samuel R. Bowman. 2019. Human vs. muppet: A conservative estimate of human performance on the GLUE benchmark. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4566–4575, Florence, Italy. Association for Computational Linguistics.
  48. 48.Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt, Clara Vania, and Samuel R. Bowman. 2021. What ingredients make for an effective crowdsourcing protocol for difficult NLU data collection tasks? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1221–1235, Online. Association for Computational Linguistics.
  49. 49.Team NLLB, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation. Meta Research.
  50. 50.Simon Ostermann, Michael Roth, and Manfred Pinkal. 2019. MCScript2.0: A machine comprehension corpus focused on script events and participants. In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019), pages 103–117, Minneapolis, Minnesota. Association for Computational Linguistics.
  51. 51.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1946–1958, Vancouver, Canada. Association for Computational Linguistics.
  52. 52.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only.
  53. 53.Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2021. UNKs everywhere: Adapting multilingual language models to new scripts. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10186–10203, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  54. 54.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  55. 55.Matthew Richardson, Christopher J.C. Burges, and Erin Renshaw. 2013. MCTest: A challenge dataset for the open-domain machine comprehension of text. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 193–203, Seattle, Washington, USA. Association for Computational Linguistics.
  56. 56.Melissa Roemmele, Cosmin Bejan, and Andrew Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In AAAI Spring Symposium - Technical Report.
  57. 57.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  58. 58.Priyanka Sen, Alham Fikri Aji, and Amir Saffari. 2022. Mintaka: A complex, natural, and multilingual dataset for end-to-end question answering. In Proceedings of the 29th International Conference on Computational Linguistics, pages 1604–1619, Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
  59. 59.Tatiana Shavrina, Alena Fenogenova, Emelyanov Anton, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, and Andrey Evlampiev. 2020. RussianSuperGLUE: A Russian language understanding evaluation benchmark. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4717–4726, Online. Association for Computational Linguistics.
  60. 60.Yuan Sun, Sisi Liu, Chaofan Chen, Zhengcuo Dan, and Xiaobing Zhao. 2021. Construction of high-quality Tibetan dataset for machine reading comprehension. In Proceedings of the 20th Chinese National Conference on Computational Linguistics, pages 208–218, Huhhot, China. Chinese Information Processing Society of China.
  61. 61.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  62. 62.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models.
  63. 63.Johannes Welbl, Nelson F. Liu, and Matt Gardner. 2017. Crowdsourcing multiple choice science questions. In Proceedings of the 3rd Workshop on Noisy User-generated Text, pages 94–106, Copenhagen, Denmark. Association for Computational Linguistics.
  64. 64.Jason Weston, Antoine Bordes, Sumit Chopra, and Tomás Mikolov. 2016. Towards ai-complete question answering: A set of prerequisite toy tasks. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings.
  65. 65.Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122, New Orleans, Louisiana. Association for Computational Linguistics.
  66. 66.Weihao Yu, Zihang Jiang, Yanfei Dong, and Jiashi Feng. 2020. Reclor: A reading comprehension dataset requiring logical reasoning. In International Conference on Learning Representations.
  67. 67.Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018. SWAG: A large-scale adversarial dataset for grounded commonsense inference. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 93–104, Brussels, Belgium. Association for Computational Linguistics.
  68. 68.Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016. Transfer learning for low-resource neural machine translation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1568–1575, Austin, Texas. Association for Computational Linguistics.

Citation

MLA
Bandarkar, L., et al. “The Belebele Benchmark: A Parallel Reading Comprehension Dataset in 122 Language Variants”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 749–75, https://doi.org/10.18653/v1/2024.acl-long.44.
APA
Bandarkar, L., Liang, D., Muller, B., Artetxe, M., Shukla, S. N., Husa, D., Goyal, N., Krishnan, A., Zettlemoyer, L., & Khabsa, M. (2024). The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 749–775. https://doi.org/10.18653/v1/2024.acl-long.44
Chicago
Bandarkar, L., D. Liang, B. Muller, et al. 2024. “The Belebele Benchmark: A Parallel Reading Comprehension Dataset in 122 Language Variants”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 749–75. https://doi.org/10.18653/v1/2024.acl-long.44.
Harvard
Bandarkar, L. et al. (2024) “The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 749–775. Available at: https://doi.org/10.18653/v1/2024.acl-long.44.
Vancouver
1. Bandarkar L, Liang D, Muller B, Artetxe M, Shukla SN, Husa D, Goyal N, Krishnan A, Zettlemoyer L, Khabsa M (2024) The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 749–775

BibTeX

@inproceedings{bandarkar-etal-2024-belebele,
    title = "The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants",
    author = "Bandarkar, Lucas  and
      Liang, Davis  and
      Muller, Benjamin  and
      Artetxe, Mikel  and
      Shukla, Satya Narayan  and
      Husa, Donald  and
      Goyal, Naman  and
      Krishnan, Abhinandan  and
      Zettlemoyer, Luke  and
      Khabsa, Madian",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.44/",
    doi = "10.18653/v1/2024.acl-long.44",
    pages = "749--775"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/