MEGA: Multilingual Evaluation of Generative AI

Kabir AhujaHarshita DiddeeRishav HadaMillicent OchiengKrithika RameshPrachi JainAkshay Uttama NambiTanuja GanuSameer SegalMohamed Ahmed

article2023EMNLP445 citations

Introduces MEGA, a comprehensive multilingual benchmark spanning 16 standard tasks across 70 typologically diverse languages to systematically compare generative large language models against prior state-of-the-art systems and identify critical performance gaps in low-resource settings.

Listen

Large Language Models (LLMs) such as GPT-3.5 and GPT-4 have demonstrated remarkable capabilities in English-language tasks, but their effectiveness across other global languages remains poorly understood. Because generative artificial intelligence is increasingly deployed in global products and public-facing services, uneven language performance risks creating severe digital inequalities. The article introduces the Multilingual Evaluation of Generative AI (MEGA) benchmark to comprehensively evaluate generative LLMs across diverse languages and standard natural language processing tasks, comparing them against specialized, fine-tuned baseline models.

The researchers assessed four major generative models (text-davinci-003, GPT-3.5-Turbo, GPT-4, and BLOOMZ) alongside strong fine-tuned baselines across 16 standard datasets covering 70 typologically diverse languages and 21 language families. The evaluation spanned five core task categories: text classification, question answering, sequence labeling, text summarization, and responsible AI metrics (gender bias and toxicity). The team tested multiple prompting approaches, including native monolingual prompting, zero-shot cross-lingual prompting with English examples, and translate-test workflows where non-English inputs are automatically translated into English prior to generation.

The findings show a substantial performance gap between English and non-English languages, which is especially severe for low-resource languages that use non-Latin scripts. While GPT-4 generally outperformed GPT-3.5 and BLOOMZ, fine-tuned models such as TULRv6 consistently surpassed generative models on structured tasks like classification, sequence labeling, and extractive question answering. For lower-resource languages like Burmese, Tamil, and Telugu, translating the input into English first (translate-test) boosted GPT-3.5-Turbo's accuracy by more than 30% relative to prompting in the native language, though a large gap compared to English baseline performance remained. In addition, tokenizer inefficiencies caused low-resource non-Latin languages to require up to ten times more tokens per word than English, which strongly correlated with lower task accuracy and dramatically higher processing costs.

These results demonstrate that organizations cannot assume off-the-shelf generative models will perform reliably or equitably across all global user bases. Relying on zero-shot prompting in low-resource native languages introduces significant operational and safety risks. Furthermore, the high tokenization burden increases financial costs for processing non-Latin scripts. The evidence indicates that common prompt engineering tricks, such as adding step-by-step reasoning or native-language prompt instructions, provide minimal benefit and can even degrade performance.

Organizations deploying multilingual AI should not rely solely on generative LLMs for critical, structured tasks in low-resource languages, but should instead favor fine-tuned specialized models or hybrid translate-test architectures where feasible. Practitioners should use English-based instruction templates and include at least four to eight few-shot examples within the native language context. In the long term, AI providers must rebalance pre-training data distributions and redesign tokenizers to handle non-Latin scripts more equitably.

Confidence in these findings is high regarding the overall performance disparities, but readers should note specific limitations. Preliminary analysis showed strong evidence of test data contamination in LLM pre-training sets, meaning the real-world performance of these models on unseen data is likely even lower than reported. Additionally, extremely under-resourced languages (such as many indigenous and African languages) remain underrepresented in standard academic benchmarks, requiring targeted localized pilot studies before making deployment decisions.

Cover for MEGA: Multilingual Evaluation of Generative AI

Abstract

Generative AI models have shown impressive performance on many Natural Language Processing tasks such as language understanding, reasoning, and language generation. An important question being asked by the AI community today is about the capabilities and limits of these models, and it is clear that evaluating generative AI is very challenging. Most studies on generative LLMs have been restricted to English and it is unclear how capable these models are at understanding and generating text in other languages. We present the first comprehensive benchmarking of generative LLMs - MEGA, which evaluates models on standard NLP benchmarks, covering 16 NLP datasets across 70 typologically diverse languages. We compare the performance of generative LLMs including Chat-GPT and GPT-4 to State of the Art (SOTA) non-autoregressive models on these tasks to determine how well generative models perform compared to the previous generation of LLMs. We present a thorough analysis of the performance of models across languages and tasks and discuss challenges in improving the performance of generative LLMs on low-resource languages. We create a framework for evaluating generative LLMs in the multilingual setting and provide directions for future progress in the field.

Table of Contents

  • 1 Introduction
  • 2 MEGA
  • 2.1 Datasets and Languages
  • 2.2 Models
  • 2.3 Evaluation Methodology
  • 2.3.1 Multilingual Prompting Strategies
  • 3 Results and Analysis
  • 3.1 Comparing different prompting strategies
  • 3.2 Comparing different models
  • 3.3 Factors Explaining Performance Trends
  • 4 Challenges in Multilingual Evaluation
  • 4.1 A Kaleidoscope of Choices.
  • 4.2 Test data contamination
  • 5 Related Work
  • 6 Conclusion
  • References
  • A Appendix
  • A.1 Tasks and Datasets
  • A.1.1 Classification
  • A.1.2 Question Answering
  • A.2 Sequences Labeling
  • A.2.1 Part of Speech Tagging
  • A.2.2 Named Entity Recognition
  • A.3 Generation
  • A.3.1 Summarization
  • A.3.2 Code-switching datasets
  • A.3.3 RAI datasets
  • A.4 Prompts
  • A.4.1 XNLI, IndicXNLI, GLUECoS NLI
  • A.4.2 PAWS-X
  • A.4.3 XCOPA
  • A.4.4 XQUAD, TyDiQA, MLQA
  • A.4.5 IndicQA
  • A.4.6 XStoryCloze
  • A.4.7 PANX
  • A.4.8 UDPOS
  • A.4.9 GLUECoS Sentiment Analysis
  • A.4.10 XLSum
  • A.4.11 Jigsaw
  • A.4.12 WinoMT
  • A.5 Handling Long Contexts
  • A.6 Factors Explaining Multilingual Capabilities of LLMs
  • A.7 Challenges in Multilingual Evaluation
  • A.8 Detailed Results

Knowls

  1. Knowl 1 — MEGA Benchmark Experimental Setup and Language Coverage

    experimental setup

    The Multilingual Evaluation of Generative AI (MEGA) benchmark is designed to evaluate the multilingual capabilities of large language models across 16 NLP datasets spanning 70 typologically diverse languages from 21 language families (with Indo-Aryan and Afro-Asiatic families forming the largest groups).

    The benchmark evaluates five core NLP task families:

    1. Classification:
      • Natural Language Inference (NLI): XNLI (15 languages), Indic-XNLI (11 Indic languages), and GLUECoS NLI (English-Hindi code-mixed).
      • Commonsense Reasoning: XCOPA (causal reasoning in 10 languages) and XStoryCloze (story ending continuation in 11 languages).
      • Paraphrase Identification: PAWS-X (7 languages).
      • Sentiment Analysis: EN-ES-CS (English-Spanish code-mixed tweets from GLUECoS).
    2. Question Answering (Span Prediction): XQuAD (11 languages), MLQA (6 languages), TyDiQA-GoldP (9 languages), and IndicQA (10 Indic languages).
    3. Sequence Labeling: Universal Dependencies POS tagging (UDPOS; 38 languages) and Named Entity Recognition (PAN-X / WikiANN; 48 languages), evaluated on the first 1,000 test examples per language.
    4. Natural Language Generation (Summarization): XL-Sum (abstractive summarization across 44 languages), evaluated on the first 1,000 test examples per language.
    5. Responsible AI (RAI): Jigsaw (toxicity classification across 6 languages) and WinoMT (gender bias in machine translation across 8 languages).

    Evaluated generative models include OpenAI's text-davinci-003 (4,096-token context limit), gpt-3.5-turbo (16k context limit), gpt-4-32k (32k context limit), and the open-source multilingual model BLOOMZ (176B parameters). These are compared against non-autoregressive multilingual fine-tuned models: TULRv6-XXL (the state of the art on XTREME), XLM-R Large, mBERT, mT5-Base, and MuRIL (specialized for Indic languages).

  2. Knowl 2 — Prompting Formulation and Multilingual Prompting Strategies

    model/method

    In MEGA, in-context learning and instruction-following for an LLM P(⋅;θ)P(\cdot; \theta) are parameterized by five components for any test query xtestx_{\text{test}}:

    1. A test input xtestx_{\text{test}}.
    2. kk few-shot exemplars {(xi,yi)}i=1k\{(x_i, y_i)\}_{i=1}^k (k=8k=8 for classification and sequence labeling; k=4k=4 for reading comprehension QA and summarization).
    3. A natural language task instruction II.
    4. A prompt template ftemp(x)f_{\text{temp}}(x) formatting dataset inputs into text.
    5. An answer verbalizer fverb(y)f_{\text{verb}}(y) mapping ground-truth label yy to a textual token.

    The composite prompt string is constructed by concatenation (∥\parallel):

    fprompt(xtest)=I∥(⨁i=1k[ftemp(xi)∥fverb(yi)])∥ftemp(xtest)f_{\text{prompt}}(x_{\text{test}}) = I \parallel \left( \bigoplus_{i=1}^k \left[ f_{\text{temp}}(x_i) \parallel f_{\text{verb}}(y_i) \right] \right) \parallel f_{\text{temp}}(x_{\text{test}})

    The model's prediction is obtained via:

    ztest=arg⁡max⁡z∈ZP(z∣fprompt(xtest);θ)z_{\text{test}} = \arg\max_{z \in \mathcal{Z}} P(z \mid f_{\text{prompt}}(x_{\text{test}}); \theta)

    where Z\mathcal{Z} denotes the model's textual output space, approximated by sampling from the predicted token distribution.

    MEGA evaluates three multilingual prompting strategies:

    • Monolingual Prompting: The kk few-shot exemplars and the test input xtestx_{\text{test}} are in the target language. Prompt templates and instructions remain in English.
    • Zero-Shot Cross-Lingual Prompting: The kk few-shot exemplars are provided in a pivot language (English), while the test query xtestx_{\text{test}} is in the target non-English language.
    • Translate-Test Prompting: The kk few-shot exemplars are in English, and the target-language test example xtestx_{\text{test}} is automatically machine-translated into English before being queried.
  3. Knowl 3 — Comparative Performance of Generative LLMs vs. Fine-Tuned Multilingual SOTA

    empirical result

    Across standard multilingual understanding, sequence labeling, and generation benchmarks, generative LLMs (GPT-3.5, GPT-4, and BLOOMZ) generally lag behind fine-tuned encoder and encoder-decoder models such as TULRv6-XXL and XLM-R Large, with the exception of commonsense reasoning tasks.

    Model XNLI PAWS-X XCOPA XStoryCloze XQuAD TyDiQA MLQA UDPOS PAN-X XL-Sum
    Acc Acc Acc Acc F1 / EM F1 / EM F1 / EM F1 F1 ROUGE-L
    mBERT 65.4 81.9 56.1 – 64.5 / 49.4 59.7 / 43.9 61.4 / 44.2 71.9 62.2 –
    mT5-Base 75.4 86.4 49.9 – 67.0 / 49.0 57.2 / 41.2 64.6 / 45.0 – 55.7 28.1
    XLM-R Large 79.2 86.4 69.2 – 76.6 / 60.8 65.1 / 45.0 71.6 / 53.2 76.2 65.2 –
    TuLRv6-XXL 88.8 93.2 82.2 – 86.0 / 72.9 84.6 / 73.8 81.0 / 63.9 83.0 84.7 –
    BLOOMZ 54.2 82.2 60.4 76.2 70.7 / 58.8 75.2 / 63.2 – – – –
    text-davinci-003 59.27 67.08 75.2 74.7 40.5 / 28.0 49.7 / 38.3 44.0 / 28.8 – – –
    gpt-3.5-turbo 62.1 70.0 79.1 87.7 60.4 / 38.2 60.1 / 38.4 56.1 / 32.8 60.2 40.3 18.8
    gpt-4-32k 75.4 73.0 89.7 96.5 68.3 / 46.6 71.5 / 50.9 67.2 / 43.3 66.6 55.5 19.7

    Key performance trends include:

    • On classification (XNLI), TULRv6-XXL (88.8% accuracy) outperforms GPT-4 (75.4%) by 13.4 points and GPT-3.5-Turbo (62.1%) by 26.7 points.
    • On sequence labeling tasks, fine-tuned baselines show substantial advantages: on PAN-X NER, TULRv6-XXL reaches 84.7 F1 compared to 55.5 for GPT-4 and 40.3 for GPT-3.5-Turbo; on UDPOS, TULRv6 reaches 83.0 F1 compared to 66.6 for GPT-4.
    • On commonsense reasoning (XCOPA, XStoryCloze), GPT-4 achieves state-of-the-art performance (89.7% on XCOPA, 96.5% on XStoryCloze), surpassing all fine-tuned baselines.
    • On Indic datasets, the specialized fine-tuned model MuRIL outperforms GPT-3.5-Turbo on IndicXNLI (76.0% vs. 50.7% accuracy) and IndicQA (47.7 vs. 38.6 F1). GPT-4 achieves 66.8% on IndicXNLI and 55.0 F1 on IndicQA.
  4. Knowl 4 — Efficacy and Disparities of Translate-Test Prompting

    empirical result

    Translating target-language test instances into English before prompting (Translate-Test) substantially improves model performance over Monolingual prompting for low-resource and non-Latin script languages, but offers minimal to negative benefits for high-resource Latin-script languages.

    Key quantitative findings on gpt-3.5-turbo:

    • Low-resource gains: Languages such as Burmese, Tamil, and Telugu achieve greater than 30%30\% relative performance gains using Translate-Test over Monolingual prompting. Malayalam, Oriya, Gujarati, Basque, Marathi, Punjabi, and Bengali show relative improvements between 15%15\% and 25%25\%.
    • High-resource plateau: High-resource European languages (German, Russian, Italian, Indonesian, French, Spanish) show near-zero or slightly negative relative gains (−1%-1\% to +3%+3\%).
    • GPT-4 behavior: While GPT-4 has stronger native non-English capabilities, Translate-Test still delivers massive gains on low-resource languages. On XStoryCloze, GPT-4's accuracy on Burmese rises from 77.6%77.6\% with Monolingual prompting to 93.2%93.2\% with Translate-Test.
    • Persistent gap to English: Even with Translate-Test, low-resource performance remains far below English performance. For example, on XNLI, Translate-Test increases gpt-3.5-turbo Urdu accuracy from 49.1%49.1\% to 54.0%54.0\%, but this remains substantially below the model's 76.2%76.2\% accuracy on native English.
  5. Knowl 5 — Impact of Tokenizer Fertility on Multilingual Performance

    empirical result

    Tokenizer fertility—defined as the average number of sub-word tokens produced per tokenized word—critically degrades downstream task performance in OpenAI LLMs (gpt-3.5-turbo and gpt-4).

    1. Fertility Disparity: In OpenAI tokenizers, high-resource Latin-script languages (English, French, German, Spanish) have fertilities between 1.0 and 1.5 tokens per word. Low-resource, non-Latin script languages (e.g., Malayalam, Tamil, Telugu, Hindi, Bengali) have fertilities between 7.0 and 10.0 tokens per word, effectively degrading tokenization to the byte level.
    2. Negative Performance Correlation: Downstream task performance exhibits statistically significant negative Pearson correlations with tokenizer fertility (p<0.05p < 0.05):
      • IndicQA: ρ=−0.960\rho = -0.960 (p=0.002p = 0.002) for GPT-3.5-Turbo; ρ=−0.856\rho = -0.856 (p=0.029p = 0.029) for GPT-4.
      • XCOPA: ρ=−0.982\rho = -0.982 (p<10−4p < 10^{-4}) for GPT-3.5-Turbo; ρ=−0.957\rho = -0.957 (p<10−4p < 10^{-4}) for GPT-4.
      • XStoryCloze: ρ=−0.745\rho = -0.745 (p=0.033p = 0.033) for GPT-3.5-Turbo; ρ=−0.918\rho = -0.918 (p=0.001p = 0.001) for GPT-4.
      • XNLI + IndicXNLI: ρ=−0.784\rho = -0.784 (p<10−4p < 10^{-4}) for GPT-3.5-Turbo; ρ=−0.803\rho = -0.803 (p<10−4p < 10^{-4}) for GPT-4.
      • XQuAD: ρ=−0.865\rho = -0.865 (p<10−3p < 10^{-3}) for GPT-3.5-Turbo; ρ=−0.818\rho = -0.818 (p=0.002p = 0.002) for GPT-4.
      • XL-Sum: ρ=−0.821\rho = -0.821 (p<10−6p < 10^{-6}) for GPT-3.5-Turbo; ρ=−0.578\rho = -0.578 (p=0.002p = 0.002) for GPT-4.
    3. Practical Impact: High fertility results in severe context length consumption and directly inflates per-sample API costs for speakers of underrepresented languages.
  6. Knowl 6 — Correlation of Pre-Training Data Size with Multilingual Performance

    empirical result

    Language representation in pre-training data correlates positively and significantly with downstream multilingual performance across task families in gpt-3.5-turbo and gpt-4.

    Evaluating the correlation between log⁡(PretrainSize)\log(\text{PretrainSize}) (using documented GPT-3 language-wise word counts as a proxy) and language-specific task performance yields positive Pearson correlation coefficients (ρ\rho):

    • XNLI: ρ=0.893\rho = 0.893 (p=4.1×10−9p = 4.1 \times 10^{-9}) for GPT-3.5-Turbo; ρ=0.836\rho = 0.836 (p=3.5×10−7p = 3.5 \times 10^{-7}) for GPT-4.
    • PAWS-X: ρ=0.850\rho = 0.850 (p=0.031p = 0.031) for GPT-3.5-Turbo; ρ=0.940\rho = 0.940 (p=0.005p = 0.005) for GPT-4.
    • XQuAD: ρ=0.782\rho = 0.782 (p=0.004p = 0.004) for GPT-3.5-Turbo; ρ=0.736\rho = 0.736 (p=0.009p = 0.009) for GPT-4.
    • XCOPA: ρ=0.700\rho = 0.700 (p=0.035p = 0.035) for GPT-3.5-Turbo; ρ=0.489\rho = 0.489 (p=0.181p = 0.181) for GPT-4.
    • MLQA: ρ=0.710\rho = 0.710 (p=0.085p = 0.085) for GPT-3.5-Turbo; ρ=0.808\rho = 0.808 (p=0.051p = 0.051) for GPT-4.
    • IndicQA: ρ=0.628\rho = 0.628 (p=0.051p = 0.051) for GPT-3.5-Turbo; ρ=0.690\rho = 0.690 (p=0.027p = 0.027) for GPT-4.

    Pre-training volume accounts for performance disparities between languages with similar tokenizer fertilities. For instance, French and Japanese share comparable tokenizer fertility in OpenAI models, yet GPT-3.5-Turbo achieves 72.1%72.1\% accuracy in French on PAWS-X versus 67.0%67.0\% in Japanese, corresponding to roughly 3.5×1093.5\times 10^9 French pre-training tokens versus 2.14×1082.14\times 10^8 Japanese tokens.

  7. Knowl 7 — Sensitivity of Multilingual LLMs to Prompt Design Choices

    empirical result

    Systematic evaluation of prompt variations across languages and datasets reveals key prompt-design behaviors for multilingual LLMs:

    1. Prompt Template Language: Translating prompt instructions/templates into the native target language reduces task accuracy compared to using English templates with target-language exemplars on text-davinci-003:
      • XNLI: 58.3% (English templates) vs. 54.4% (Native templates)
      • Indic-XNLI: 49.6% vs. 38.7%
      • PAWS-X: 67.1% vs. 64.2%
      • XCOPA: 77.6% vs. 73.1%
    2. Few-Shot Exemplar Scaling (kk): Performance on XNLI and XCOPA increases sharply when transitioning from k=0k=0 to k=2–4k=2\text{--}4, but plateaus for k≥8k \ge 8 across most languages. In contrast, very low-resource languages (e.g., Haitian Creole on XCOPA) continue to exhibit accuracy gains up to k=16k=16.
    3. Language-Specific Prompt Tuning: Selecting prompt templates on target-language validation data rather than English validation data improves accuracy for low-resource languages with sufficient validation size (e.g., Haitian Creole on XCOPA improved from 72.0%72.0\% to 75.6%75.6\%), but degrades performance when target-language validation sets are small (e.g., Tamil on XCOPA with N=100N=100).
    4. English Explanations: Incorporating English intermediate explanations (Explain-then-Predict) into few-shot exemplars does not systematically improve multilingual accuracy on XStoryCloze and XCOPA, and can decrease accuracy by 3–5%3\text{--}5\% on low-resource languages (e.g., Haitian Creole and Tamil). Error inspection indicates the model unpromptedly translates the non-English premise into English before generating the explanation.
  8. Knowl 8 — Retrieve-then-Prompt QA Strategy and Low-Resource Retrieval Failure

    model/method

    For reading comprehension QA on models with small context windows (text-davinci-003 with a 4,096-token limit), high tokenizer fertility prevents fitting full context paragraphs and few-shot examples into the prompt. A retrieve-then-prompt pipeline addresses this constraint:

    1. Few-Shot Context Truncation: For in-context exemplars, only the single sentence containing the ground-truth answer span is retained as context.
    2. Test Context Chunking and Semantic Retrieval: The test context document is split into chunks of maximum length 100 tokens. Each chunk is embedded using text-embedding-ada-002. The question is embedded with the same model, and the chunk with highest cosine similarity to the question is retrieved and inserted into the test prompt.

    Multilingual Retrieval Failure Mode: Retrieval accuracy—defined as the proportion of instances where the retrieved 100-token chunk contains the gold answer span—collapses on non-Latin, low-resource languages on TyDiQA:

    • English (en): 85.8%
    • Swahili (sw): 76.0%
    • Finnish (fi): 75.6%
    • Indonesian (id): 68.0%
    • Arabic (ar): 49.2%
    • Korean (ko): 45.3%
    • Russian (ru): 42.1%
    • Bengali (bn): 14.1%
    • Telugu (te): 5.6%

    This retrieval bottleneck causes text-davinci-003 to score only 8.45 F1 on IndicQA and 49.8 F1 on TyDiQA, whereas models that process unchunked contexts (gpt-3.5-turbo and gpt-4) achieve 38.6 F1 and 55.0 F1 on IndicQA, respectively.

  9. Knowl 9 — Test Data Contamination Analysis in Multilingual Evaluation

    empirical result

    To evaluate whether multilingual LLM performance is influenced by memorization of evaluation benchmarks, GPT-4 was assessed across three contamination indicators on MEGA datasets:

    1. Card Fill: Prompting GPT-4 to generate structured dataset cards (supported languages, task definition, input-output schema). Performance is rated as Full (all metadata correct), Partial (partially correct), or None.
    2. Data Accessibility without Download (Data Acc. w/o Down.): Verifying whether test splits are directly viewable on web pages or online dataset viewers without downloading archive files.
    3. Release Date: Checking whether dataset release preceded the September 2021 pre-training cutoff.
    Dataset Card Fill Data Acc. w/o Down. Release Date
    XNLI Full Yes September 2019
    Indic-XNLI Full Yes April 2022
    PAWS-X Full Yes August 2019
    XCOPA Partial Yes April 2020
    XStoryCloze Partial No May 2023
    XQuAD Full Yes October 2019
    MLQA Full Yes October 2019
    TyDiQA-GoldP Full Yes February 2020
    IndicQA Partial Yes September 2022
    PAN-X Full Yes July 2017
    UDPOS Full Yes March 2020
    XLSum Partial Yes June 2021
    Jigsaw None No February 2020
    GLUECos NLI None No June 2020
    EN-ES-CS None No May 2016

    Key takeaways:

    • GPT-4 achieves Full card completion on 8 benchmarks and Partial completion on 4 benchmarks; 12 of 15 test sets are viewable online without downloading.
    • Jigsaw, GLUECoS NLI, and EN-ES-CS show no evidence of contamination (Card Fill: None, Data Acc w/o Down: No).
    • Because LLMs underperform on low-resource non-English languages despite high likelihood of contamination across standard multilingual datasets, the true generalization disparity between English and low-resource languages is likely larger than benchmark scores indicate.
  10. Knowl 10 — Multilingual Gender Bias and Toxicity Evaluation in Generative LLMs

    empirical result

    Evaluation of generative LLMs on Responsible AI dimensions demonstrates pronounced multilingual disparities in gender bias and toxicity detection:

    1. Gender Bias in Translation (WinoMT): Evaluating zero-shot translation of 3,888 coreference sentences across 8 target languages using gender accuracy (Acc), gender disparity ΔG=∣F1masc−F1fem∣\Delta G = |\text{F1}_{\text{masc}} - \text{F1}_{\text{fem}}|, and stereotypical bias disparity ΔS=∣F1pro−F1anti∣\Delta S = |\text{F1}_{\text{pro}} - \text{F1}_{\text{anti}}|:
      • gpt-3.5-turbo overall gender accuracy ranges from 41.0% (Russian) to 61.1% (Arabic).
      • gpt-3.5-turbo exhibits strong pro-stereotypical bias (ΔS\Delta S), reaching 40.8 on Hebrew, 27.9 on Arabic, 26.7 on Italian, 26.2 on Spanish, and 26.1 on French.
      • BLOOMZ exhibits severe masculine skew ΔG\Delta G (e.g., 56.2 on German) and produces zero precision for masculine predictions on Russian, leading to invalid ΔG\Delta G.
    2. Multilingual Toxicity Classification (Jigsaw): Evaluating zero-shot crosslingual prompting and translate-test on 6 test languages (Turkish, Portuguese, Russian, Spanish, Italian, French):
      • gpt-3.5-turbo achieves an average accuracy of 78.79% under crosslingual prompting (ranging from 73.64% on French to 85.65% on Turkish) and 75.05% under Translate-Test.
      • text-davinci-003 achieves 81.50% average crosslingual accuracy (ranging from 74.55% on French to 93.55% on Turkish) and 79.97% under Translate-Test.
      • Both models trail PaLM-2 (10-shot monolingual average of 91.65%, reaching 94.34% on Turkish and 94.25% on Russian).

Coverage note — No substantial contributed material was omitted. Full language-by-language raw data appendices for individual datasets are represented through their aggregated benchmark tables, correlation statistics, and representative language case studies.

References

  1. 1.David Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba Alabi, Shamsuddeen Muhammad, Peter Nabende, Cheikh M. Bamba Dione, Andiswa Bukula, Rooweither Mabuya, Bonaventure F. P. Dossou, Blessing Sibanda, Happy Buzaaba, Jonathan Mukiibi, Godson Kalipe, Derguene Mbaye, Amelia Taylor, Fatoumata Kabore, Chris Chinenye Emezue, Anuoluwapo Aremu, Perez Ogayo, Catherine Gitau, Edwin Munkoh-Buabeng, Victoire Memdjokam Koagne, Allahsera Auguste Tapo, Tebogo Macucwa, Vukosi Marivate, Mboning Tchiaze Elvis, Tajuddeen Gwada-be, Tosin Adewumi, Orevaoghene Ahia, Joyce Nakatumba-Nabende, Neo Lerato Mokono, Ignatius Ezeani, Chiamaka Chukwuneke, Mofetoluwa Oluwaseun Adeyemi, Gilles Quentin Hacheme, Idris Abdulmumin, Odunayo Ogundepo, Oreen Yousuf, Tatiana Moteu, and Dietrich Klakow. 2022. MasakhaNER 2.0: Africa-centric transfer learning for named entity recognition. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4488–4508, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  2. 2.Divyanshu Aggarwal, Vivek Gupta, and Anoop Kunchukuttan. 2022. Indicxnli: Evaluating multilingual inference for indian languages. arXiv preprint arXiv:2204.08776.
  3. 3.Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David R. Mortensen, Noah A. Smith, and Yulia Tsvetkov. 2023. Do all languages cost the same? tokenization in the era of commercial language models. ArXiv, abs/2305.13707.
  4. 4.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637.
  5. 5.Akari Asai, Sneha Kudugunta, Xinyan Velocity Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajishirzi. 2023. Buffet: Benchmarking large language models for few-shot cross-lingual transfer. arXiv cs.CL 2305.14857.
  6. 6.Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Fries, Maged Alshaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Dragomir Radev, Mike Tian-jian Jiang, and Alexander Rush. 2022. PromptSource: An integrated development environment and repository for natural language prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 93–104, Dublin, Ireland. Association for Computational Linguistics.
  7. 7.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, et al. 2023. A multi-task, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity. arXiv preprint arXiv:2302.04023.
  8. 8.Damian Blasi, Antonios Anastasopoulos, and Graham Neubig. 2022. Systematic inequalities in language technology performance across the world’s languages. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5486–5505, Dublin, Ireland. Association for Computational Linguistics.
  9. 9.Terra Blevins, Hila Gonen, and Luke Zettlemoyer. 2022. Prompting language models for linguistic structure. arXiv cs.CL 2211.07830.
  10. 10.Terra Blevins and Luke Zettlemoyer. 2022. Language contamination helps explains the cross-lingual capabilities of English pretrained models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3563–3574, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  11. 11.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. CoRR, abs/2005.14165.
  12. 12.Monojit Choudhury and Amit Deshpande. 2021. How linguistically fair are multilingual pre-trained language models? In AAAI-21. AAAI, AAAI.
  13. 13.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  14. 14.Jonathan H Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics, 8:454–470.
  15. 15.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451.
  16. 16.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of EMNLP 2018, pages 2475–2485.
  17. 17.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of naacL-HLT, pages 4171–4186.
  18. 18.Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M Khapra, Anoop Kunchukuttan, and Pratyush Kumar. 2022. Indicxtreme: A multi-task benchmark for evaluating indic languages. arXiv preprint arXiv:2212.05409.
  19. 19.A. Seza Dogruöz, Sunayana Sitaram, Barbara E. Bullock, and Almeida Jacqueline Toribio. 2021. A survey of code-switching: Linguistic and social perspectives for language technologies. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1654–1666, Online. Association for Computational Linguistics.
  20. 20.Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M Sohel Rahman, and Rifat Shahriyar. 2021a. Xlsum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703.
  21. 21.Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021b. XLSum: Large-scale multilingual abstractive summarization for 44 languages. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4693–4703, Online. Association for Computational Linguistics.
  22. 22.Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are gpt models at machine translation? a comprehensive evaluation. arXiv preprint arXiv:2302.09210.
  23. 23.Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In International Conference on Machine Learning, pages 4411–4421. PMLR.
  24. 24.Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6282–6293, Online. Association for Computational Linguistics.
  25. 25.Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, et al. 2021. Muril: Multilingual representations for indian languages. arXiv preprint arXiv:2103.10730.
  26. 26.Simran Khanuja, Sandipan Dandapat, Sunayana Sitaram, and Monojit Choudhury. 2020a. A new dataset for natural language inference from code-mixed conversations. In Proceedings of the The 4th Workshop on Computational Approaches to Code Switching, pages 9–16.
  27. 27.Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, and Monojit Choudhury. 2020b. Gluecos: An evaluation benchmark for code-switched nlp. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3575–3585.
  28. 28.Ian Kivlichan, Jeffrey Sorensen, Julia Elliott, Lucy Vasserman, Martin Görner, and Phil Culliton. 2020. Jigsaw multilingual toxic comment classification.
  29. 29.Viet Dac Lai, Nghia Trung Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023. Chatgpt beyond english: Towards a comprehensive evaluation of large language models in multilingual learning.
  30. 30.Anne Lauscher, Vinit Ravishankar, Ivan Vulic, and Goran Glavaš. 2020. From zero to hero: On the limitations of zero-shot language transfer with multilingual Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4483–4499, Online. Association for Computational Linguistics.
  31. 31.Patrick Lewis, Barlas Oguz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2020. Mlqa: Evaluating cross-lingual extractive question answering. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7315–7330.
  32. 32.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  33. 33.Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, et al. 2020. Xglue: A new benchmark dataset for cross-lingual pre-training, understanding and generation. arXiv preprint arXiv:2004.01401.
  34. 34.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2022a. Few-shot learning with multilingual generative language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9019–9052, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  35. 35.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2022b. Few-shot learning with multilingual generative language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9019–9052.
  36. 36.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022. What makes good in-context examples for GPT-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114, Dublin, Ireland and Online. Association for Computational Linguistics.
  37. 37.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  38. 38.Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. Cross-task generalization via natural language crowdsourcing instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3470–3487, Dublin, Ireland. Association for Computational Linguistics.
  39. 39.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 839–849, San Diego, California. Association for Computational Linguistics.
  40. 40.Nasrin Mostafazadeh, Michael Roth, Annie Louis, Nathanael Chambers, and James Allen. 2017. Lsdsem 2017 shared task: The story cloze test. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics, pages 46–51.
  41. 41.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2022. Crosslingual generalization through multitask finetuning.
  42. 42.Akshay Nambi, Vaibhav Balloli, Mercy Ranjit, Tanuja Ganu, Kabir Ahuja, Sunayana Sitaram, and Kalika Bali. 2023. Breaking language barriers with a leap: Learning strategies for polyglot llms. arXiv cs.CL 2305.17740.
  43. 43.Joakim Nivre, Mitchell Abrams, Željko Agic, Lars Ahrenberg, Lene Antonsen, Maria Jesus Aranzabe, Gashaw Arutie, Masayuki Asahara, Luma Ateyah, Mohammed Attia, et al. 2018. Universal dependencies 2.2.
  44. 44.Maxwell I. Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2021. Show your work: Scratchpads for intermediate computation with language models. CoRR, abs/2112.00114.
  45. 45.OpenAI. 2023. Gpt4 technical report.
  46. 46.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  47. 47.Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. Cross-lingual name tagging and linking for 282 languages. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1946–1958.
  48. 48.Barun Patra, Saksham Singhal, Shaohan Huang, Zewen Chi, Li Dong, Furu Wei, Vishrav Chaudhary, and Xia Song. 2022. Beyond english-centric bitexts for better multilingual language representation learning. arXiv preprint arXiv:2210.14867.
  49. 49.Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulic, and Anna Korhonen. 2020. Xcopa: A multilingual dataset for causal commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2362–2376.
  50. 50.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016a. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  51. 51.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016b. Squad: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392.
  52. 52.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In AAAI spring symposium: logical formalizations of commonsense reasoning, pages 90–95.
  53. 53.Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, et al. 2021. Xtreme-r: Towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10215–10245.
  54. 54.Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender bias in coreference resolution. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 8–14, New Orleans, Louisiana. Association for Computational Linguistics.
  55. 55.Phillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder, and Iryna Gurevych. 2021. How good is your tokenizer? on the monolingual performance of multilingual language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3118–3135, Online. Association for Computational Linguistics.
  56. 56.Oscar Sainz, Jon Ander Campos, Iker García-Ferrero, Julen Etxaniz, and Eneko Agirre. 2023. Did chatgpt cheat on your test?
  57. 57.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  58. 58.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools.
  59. 59.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2022. Language models are multilingual chain-of-thought reasoners. CoRR, abs/2210.03057.
  60. 60.Andy Shih, Dorsa Sadigh, and Stefano Ermon. 2023. Long horizon temperature scaling.
  61. 61.Sunayana Sitaram, Khyathi Raghavi Chandu, Sai Krishna Rallabandi, and Alan W Black. 2019. A survey of code-switched speech and language processing. arXiv preprint arXiv:1904.00784.
  62. 62.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ambrose Slone, Ameet Rahane, Anantharaman S. Iyer, Anders Andreassen, Andrea Madotto, Andrea Santilli, Andreas Stuhlmüller, Andrew Dai, Andrew La, Andrew Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karakaş, B. Ryan Roberts, Bao Sheng Loe, Barret Zoph, Bartłomiej Bojanowski, Batuhan Özyurt, Behnam Hedayatnia, Behnam Neyshabur, Benjamin Inden, Benno Stein, Berk Ekmekci, Bill Yuchen Lin, Blake Howald, Bryan Orinion, Cameron Diao, Cameron Dour, Catherine Stinson, Cedrick Argueta, César Ferri Ramírez, Chandan Singh, Charles Rathkopf, Chenlin Meng, Chitta Baral, Chiyu Wu, Chris Callison-Burch, Chris Waites, Christian Voigt, Christopher D. Manning, Christopher Potts, Cindy Ramirez, Clara E. Rivera, Clemencia Siro, Colin Raffel, Courtney Ashcraft, Cristina Garbacea, Damien Sileo, Dan Garrette, Dan Hendrycks, Dan Kilman, Dan Roth, Daniel Freeman, Daniel Khashabi, Daniel Levy, Daniel Moseguí González, Danielle Perszyk, Danny Hernandez, Danqi Chen, Daphne Ippolito, Dar Gilboa, David Dohan, David Drakard, David Jurgens, Debajyoti Datta, Deep Ganguli, Denis Emelin, Denis Kleyko, Deniz Yuret, Derek Chen, Derek Tam, Dieuwke Hupkes, Diganta Misra, Dilyar Buzan, Dimitri Coelho Mollo, Diyi Yang, Dong-Ho Lee, Dylan Schrader, Ekaterina Shutova, Ekin Dogus Cubuk, Elad Segal, Eleanor Hagerman, Elizabeth Barnes, Elizabeth Donoway, Ellie Pavlick, Emanuele Rodola, Emma Lam, Eric Chu, Eric Tang, Erkut Erdem, Ernie Chang, Ethan A. Chi, Ethan Dyer, Ethan Jerzak, Ethan Kim, Eunice Engefu Manyasi, Evgenii Zheltonozhskii, Fanyue Xia, Fatemeh Siar, Fernando Martínez-Plumed, Francesca Happé, Francois Chollet, Frieda Rong, Gaurav Mishra, Genta Indra Winata, Gerard de Melo, Germán Kruszewski, Giambattista Parascandolo, Giorgio Mariani, Gloria Wang, Gonzalo Jaimovitch-López, Gregor Betz, Guy Gur-Ari, Hana Galijasevic, Hannah Kim, Hannah Rashkin, Hannaneh Hajishirzi, Harsh Mehta, Hayden Bogar, Henry Shevlin, Hinrich Schütze, Hiromu Yakura, Hongming Zhang, Hugh Mee Wong, Ian Ng, Isaac Noble, Jaap Jumelet, Jack Geissinger, Jackson Kernion, Jacob Hilton, Jaehoon Lee, Jaime Fernández Fisac, James B. Simon, James Koppel, James Zheng, James Zou, Jan Kocon, Jana Thompson, Janelle Wingfield, Jared Kaplan, Jarema Radom, Jascha Sohl-Dickstein, Jason Phang, Jason Wei, Jason Yosinski, Jekaterina Novikova, Jelle Bosscher, Jennifer Marsh, Jeremy Kim, Jeroen Taal, Jesse Engel, Jesujoba Alabi, Jiacheng Xu, Jiaming Song, Jillian Tang, Joan Waweru, John Burden, John Miller, John U. Balis, Jonathan Batchelder, Jonathan Berant, Jörg Frohberg, Jos Rozen, Jose Hernandez-Orallo, Joseph Boudeman, Joseph Guerr, Joseph Jones, Joshua B. Tenenbaum, Joshua S. Rule, Joyce Chua, Kamil Kanclerz, Karen Livescu, Karl Krauth, Karthik Gopalakrishnan, Katerina Ignatyeva, Katja Markert, Kaustubh D. Dhole, Kevin Gimpel, Kevin Omondi, Kory Mathewson, Kristen Chiafullo, Ksenia Shkaruta, Kumar Shridhar, Kyle McDonell, Kyle Richardson, Laria Reynolds, Leo Gao, Li Zhang, Liam Dugan, Lianhui Qin, Lidia Contreras-Ochando, Louis-Philippe Morency, Luca Moschella, Lucas Lam, Lucy Noble, Ludwig Schmidt, Luheng He, Luis Oliveros Colón, Luke Metz, Lütfi Kerem Senel, Maarten Bosma, Maarten Sap, Maartje ter Hoeve, Maheen Farooqi, Manaal Faruqui, Mantas Mazeika, Marco Baturan, Marco Marelli, Marco Maru, Maria Jose Ramírez Quintana, Marie Tolkiehn, Mario Giulianelli, Martha Lewis, Martin Potthast, Matthew L. Leavitt, Matthias Hagen, Mátyás Schubert, Medina Orduna Baitemirova, Melody Arnaud, Melvin McElrath, Michael A. Yee, Michael Cohen, Michael Gu, Michael Ivanitskiy, Michael Starritt, Michael Strube, Michał Swędrowski, Michele Bevilacqua, Michihiro Yasunaga, Mihir Kale, Mike Cain, Mimee Xu, Mirac Suzgun, Mitch Walker, Mo Tiwari, Mohit Bansal, Moin Aminnaseri, Mor Geva, Mozhdeh Gheini, Mukund Varma T, Nanyun Peng, Nathan A. Chi, Nayeon Lee, Neta Gur-Ari Krakover, Nicholas Cameron, Nicholas Roberts, Nick Doiron, Nicole Martinez, Nikita Nangia, Niklas Deckers, Niklas Muennighoff, Nitish Shirish Keskar, Niveditha S. Iyer, Noah Constant, Noah Fiedel, Nuan Wen, Oliver Zhang, Omar Agha, Omar Elbaghdadi, Omer Levy, Owain Evans, Pablo Antonio Moreno Casares, Parth Doshi, Pascale Fung, Paul Vicol, Pegah Alipoormolbas hi, Peiyuan Liao, Percy Liang, Peter Chang, Peter Eckersley, Phu Mon Htut, Pinyu Hwang, Piotr Miłkowski, Piyush Patil, Pouya Pezeshkpour, Priti Oli, Qiaozhu Mei, Qing Lyu, Qinlang Chen, Rabin Banjade, Rachel Etta Rudolph, Raefer Gabriel, Rahel Habacker, Ramon Risco, Raphaël Millière, Rhythm Garg, Richard Barnes, Rif A. Saurous, Riku Arakawa, Robbe Raymaekers, Robert Frank, Rohan Sikand, Roman Novak, Roman Sitelew, Ronan LeBras, Rosanne Liu, Rowan Jacobs, Rui Zhang, Ruslan Salakhutdinov, Ryan Chi, Ryan Lee, Ryan Stovall, Ryan Teehan, Rylan Yang, Sahib Singh, Saif M. Mohammad, Sajant Anand, Sam Dillavou, Sam Shleifer, Sam Wiseman, Samuel Gruetter, Samuel R. Bowman, Samuel S. Schoenholz, Sanghyun Han, Sanjeev Kwatra, Sarah A. Rous, Sarik Ghazarian, Sayan Ghosh, Sean Casey, Sebastian Bischoff, Sebastian Gehrmann, Sebastian Schuster, Sepideh Sadeghi, Shadi Hamdan, Sharon Zhou, Shashank Srivastava, Sherry Shi, Shikhar Singh, Shima Asaadi, Shixiang Shane Gu, Shubh Pachchigar, Shubham Toshniwal, Shyam Upadhyay, Shyamolima Debnath, Siamak Shakeri, Simon Thormeyer, Simone Melzi, Siva Reddy, Sneha Priscilla Makini, Soo-Hwan Lee, Spencer Torene, Sriharsha Hatwar, Stanislas Dehaene, Stefan Divic, Stefano Ermon, Stella Biderman, Stephanie Lin, Stephen Prasad, Steven T. Piantadosi, Stuart M. Shieber, Summer Misherghi, Svetlana Kiritchenko, Swaroop Mishra, Tal Linzen, Tal Schuster, Tao Li, Tao Yu, Tariq Ali, Tatsu Hashimoto, Te-Lin Wu, Théo Desbordes, Theodore Rothschild, Thomas Phan, Tianle Wang, Tiberius Nkinyili, Timo Schick, Timofei Kornev, Titus Tunduny, Tobias Gerstenberg, Trenton Chang, Trishala Neeraj, Tushar Khot, Tyler Shultz, Uri Shaham, Vedant Misra, Vera Demberg, Victoria Nyamai, Vikas Raunak, Vinay Ramasesh, Vinay Uday Prabhu, Vishakh Padmakumar, Vivek Srikumar, William Fedus, William Saunders, William Zhang, Wout Vossen, Xiang Ren, Xiaoyu Tong, Xinran Zhao, Xinyi Wu, Xudong Shen, Yadollah Yaghoobzadeh, Yair Lakretz, Yangqiu Song, Yasaman Bahri, Yejin Choi, Yichi Yang, Yiding Hao, Yifu Chen, Yonatan Belinkov, Yu Hou, Yufang Hou, Yuntao Bai, Zachary Seid, Zhuoye Zhao, Zijian Wang, Zijie J. Wang, Zirui Wang, and Ziyi Wu. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.
  63. 63.Gabriel Stanovsky, Noah A Smith, and Luke Zettlemoyer. 2019. Evaluating gender bias in machine translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1679–1684.
  64. 64.Mike Thelwall. 2017. The heart and soul of the web? sentiment strength detection in the social web with sentistrength. Cyberemotions: Collective emotions in cyberspace, pages 119–134.
  65. 65.David Vilares, Miguel A Alonso, and Carlos Gómez-Rodríguez. 2016. En-es-cs: An english-spanish code-switching twitter corpus for multilingual sentiment analysis. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 4149–4153.
  66. 66.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. Glue: A multi-task benchmark and analysis platform for natural language understanding. EMNLP 2018, page 353.
  67. 67.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen. 2022. Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5085–5109, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  68. 68.Tom Warren. 2023. Microsoft’s chatgpt event live blog.
  69. 69.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. Finetuned language models are zero-shot learners. CoRR, abs/2109.01652.
  70. 70.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-thought prompting elicits reasoning in large language models. arXiv 2201.11903 cs.CL.
  71. 71.Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, et al. 2020. Indonlu: Benchmark and resources for evaluating indonesian natural language understanding. arXiv preprint arXiv:2009.05387.
  72. 72.Shijie Wu and Mark Dredze. 2020. Are all languages created equal in multilingual BERT? In Proceedings of the 5th Workshop on Representation Learning for NLP, pages 120–130, Online. Association for Computational Linguistics.
  73. 73.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mt5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498.
  74. 74.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019a. PAWS-X: A cross-lingual adversarial dataset for paraphrase identification. In Proceedings of EMNLP 2019, pages 3685–3690.
  75. 75.Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019b. Paws-x: A cross-lingual adversarial dataset for paraphrase identification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3687–3692.
  76. 76.Xi Ye and Greg Durrett. 2022a. Can explanations be useful for calibrating black box models? In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6199–6212, Dublin, Ireland. Association for Computational Linguistics.
  77. 77.Xi Ye and Greg Durrett. 2022b. The unreliability of explanations in few-shot prompting for textual reasoning. In Advances in Neural Information Processing Systems, volume 35, pages 30378–30392. Curran Associates, Inc.
  78. 78.Daniel Zeman, Joakim Nivre, Mitchell Abrams, Elia Ackermann, Noëmi Aepli, Hamid Aghaei, and R Ziane. 2020. Universal dependencies 2.5. LINDAT/CLARIAHCZ digital library at the Institute of Formal and Applied Linguistics (UFAL), Faculty of Mathematics and Physics, Charles University. url: http://hdl.handle.net/11234/1-3226.
  79. 79.Yuan Zhang, Jason Baldridge, and Luheng He. 2019. Paws: Paraphrase adversaries from word scrambling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1298–1308.
  80. 80.Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 15–20, New Orleans, Louisiana. Association for Computational Linguistics.
  81. 81.Mengjie Zhao and Hinrich Schütze. 2021. Discrete and soft prompting for multilingual models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8547–8555, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  82. 82.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pages 12697–12706. PMLR.

Citation

MLA
Ahuja, K., et al. “MEGA: Multilingual Evaluation of Generative AI”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 4232–67, https://doi.org/10.18653/v1/2023.emnlp-main.258.
APA
Ahuja, K., Diddee, H., Hada, R., Ochieng, M., Ramesh, K., Jain, P., Nambi, A., Ganu, T., Segal, S., Ahmed, M., Bali, K., & Sitaram, S. (2023). MEGA: Multilingual Evaluation of Generative AI. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4232–4267. https://doi.org/10.18653/v1/2023.emnlp-main.258
Chicago
Ahuja, K., H. Diddee, R. Hada, et al. 2023. “MEGA: Multilingual Evaluation of Generative AI”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4232–67. https://doi.org/10.18653/v1/2023.emnlp-main.258.
Harvard
Ahuja, K. et al. (2023) “MEGA: Multilingual Evaluation of Generative AI”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4232–4267. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.258.
Vancouver
1. Ahuja K, Diddee H, Hada R, et al (2023) MEGA: Multilingual Evaluation of Generative AI. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4232–4267

BibTeX

@inproceedings{ahuja-etal-2023-mega,
    title = "{MEGA}: Multilingual Evaluation of Generative {AI}",
    author = "Ahuja, Kabir  and
      Diddee, Harshita  and
      Hada, Rishav  and
      Ochieng, Millicent  and
      Ramesh, Krithika  and
      Jain, Prachi  and
      Nambi, Akshay  and
      Ganu, Tanuja  and
      Segal, Sameer  and
      Ahmed, Mohamed  and
      Bali, Kalika  and
      Sitaram, Sunayana",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.258/",
    doi = "10.18653/v1/2023.emnlp-main.258",
    pages = "4232--4267"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/