GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP

Md. Tawkat Islam KhondakerAbdul WaheedEl Moatez Billah NagoudiMuhammad Abdul-Mageed

article2023EMNLP108 citations

Demonstrates through a large-scale evaluation across 44 tasks and over 60 datasets that ChatGPT and GPT-4 consistently lag behind smaller, fine-tuned dedicated models on Arabic natural language processing, particularly when handling regional dialectal varieties compared to Modern Standard Arabic.

Listen

Large language models such as ChatGPT have demonstrated strong capabilities across standard English-language benchmarks, driving rapid adoption across global industries. However, their effectiveness in linguistically rich and complex languages like Arabic—spoken by more than 450 million people across diverse regional dialects—remains largely unverified at scale. The article systematically evaluates ChatGPT across a wide range of Arabic natural language understanding and generation tasks to determine whether general-purpose commercial models can reliably serve Arabic-speaking users without task-specific customization.

The article conducts an extensive empirical assessment across 44 distinct Arabic language understanding and text generation tasks spanning more than 60 datasets. The authors test ChatGPT against an instruction-tuned multilingual baseline (BLOOMZ) and two dedicated, smaller Arabic models that were specifically finetuned for Arabic understanding and generation (MARBERTV2 and AraT5). Evaluations were run across zero-shot and few-shot prompt setups using English instruction templates, complemented by targeted evaluations of country-level dialects and qualitative assessments conducted by both native human annotators and GPT-4.

The evaluation reveals several critical findings. First, despite its massive scale and general multilingual pretraining, ChatGPT is consistently outperformed by significantly smaller, Arabic-dedicated finetuned models across almost all understanding and generation benchmarks. On overall language understanding, the dedicated finetuned model achieves an aggregate macro-F1 score of approximately 69% to 71%, whereas ChatGPT peaks at roughly 51% even with ten-shot prompting. Second, both ChatGPT and GPT-4 demonstrate a substantial performance drop when processing Dialectal Arabic compared to Modern Standard Arabic, with GPT-4 outperforming ChatGPT by an average of about 19% on standard Arabic and 10% on regional dialects. Third, in text generation tasks such as machine translation, summarization, and grammatical error correction, task-dedicated models consistently achieve higher accuracy and lower error rates, although ChatGPT outperforms other models on code-switched inputs involving mixtures of Arabic and English or French. Finally, human and automated quality assessments reveal that GPT-4 ratings agree with human evaluators approximately 71.5% of the time.

These findings indicate that relying on general-purpose, out-of-the-box language models for Arabic applications introduces notable operational, quality, and compliance risks. Deploying standard commercial models in production without adaptation could lead to degraded user experiences, particularly for populations communicating in regional dialects rather than formal modern Arabic. Additionally, the analysis revealed that ChatGPT frequently exhibits high false-positive rates on content moderation tasks, often incorrectly classifying benign dialectal text as toxic. This tendency poses a risk of unnecessary content censorship or flawed moderation workflows if deployed without safeguards.

Organizations seeking to implement Arabic artificial intelligence solutions should prioritize specialized, finetuned models or explore hybrid architectures rather than relying entirely on generic foundation models. When commercial models like GPT-4 are utilized, practitioners should establish robust evaluation pipelines and restrict deployment to Modern Standard Arabic workflows until dialect handling improves. The findings are based on controlled test samples and fixed model snapshots from early 2023, meaning performance may evolve as commercial models update; nevertheless, the results provide high confidence that dedicated domain adaptation remains essential for reliable Arabic natural language processing.

arXiv: 2305.14976
Cover for GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP

Abstract

ChatGPT’s emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks. However, the model’s efficacy across diverse linguistic contexts remains largely uncharted territory. This work aims to bridge this knowledge gap, with a primary focus on assessing ChatGPT’s capabilities on Arabic languages and dialectal varieties. Our comprehensive study conducts a large-scale automated and human evaluation of ChatGPT, encompassing 44 distinct language understanding and generation tasks on over 60 different datasets. To our knowledge, this marks the first extensive performance analysis of ChatGPT’s deployment in Arabic NLP. Our findings indicate that, despite its remarkable performance in English, ChatGPT is consistently surpassed by smaller models that have undergone finetuning on Arabic. We further undertake a meticulous comparison of ChatGPT and GPT-4’s Modern Standard Arabic (MSA) and Dialectal Arabic (DA), unveiling the relative shortcomings of both models in handling Arabic dialects compared to MSA. Although we further explore and confirm the utility of employing GPT-4 as a potential alternative for human evaluation, our work adds to a growing body of research underscoring the limitations of ChatGPT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Datasets
  • 4 Prompt Design
  • 5 Experiments
  • 6 Evaluation on NLU Tasks
  • 7 Evaluation on NLG Tasks
  • 8 Performance on Dialectal Arabic
  • 9 Human Evaluation
  • 10 Using GPT-4 to Evaluate Responses
  • 11 Conclusion
  • 12 Limitations
  • 13 Ethics Statement
  • Acknowledgments
  • References
  • A Literature Review
  • A.1 Machine Translation
  • A.2 QA
  • A.3 Text Classification
  • A.4 Multilinguality
  • B Dataset
  • B.1 NLU Tasks
  • B.2 NLG Tasks
  • C Anisotropy in BERT Models
  • D ChatGPT Exhibits False-Toxicity
  • E Prompt Template for GPT-4 Evaluation
  • F Class Distribution for Few-shot Examples
  • G Human Evaluation Framework
  • G.1 Task Description
  • H Examples of Prompt
  • I Examples of Models' Response
  • J Examples of GPT-4 Evaluation
  • K GPT-4 Explanation on Models' Evaluation
  • L Human Evaluation

Knowls

  1. Knowl 1 — The benchmark spans 44 Arabic NLP tasks

    data/table

    GPTAraEval evaluates Arabic natural-language understanding (NLU) and generation (NLG). Its NLU portion draws on 48 datasets from the ORCA benchmark and covers 20 tasks: Arabic natural-language inference, claim prediction, three levels of dialect identification, machine-generated-text detection, paraphrase detection, sentiment and emotion, nine social-meaning tasks, stance detection, and word-sense disambiguation. Its NLG portion contains 24 tasks from 23 publicly available datasets, organized into 13 clusters: code-switched translation, diacritization, dialect translation, grammatical-error correction, machine translation, news-title generation, dialectal dialogue generation, paraphrasing, question answering, question generation, text rewriting, transliteration, and summarization. Together the authors describe the evaluation as covering 44 tasks. The benchmark therefore compares models across both standardized Arabic and dialectal varieties, and across classification and generation.

  2. Knowl 2 — Evaluation protocol and prompt design

    experimental setup

    The study evaluates ChatGPT (the March 1, 2023 snapshot of gpt-3.5-turbo) in 0-, 3-, 5-, and 10-shot settings, with temperature 0. For each dataset, the authors randomly sample 200 test examples. Few-shot examples are randomly drawn from the training set, with nested sampling so that examples used at a lower shot count also appear at higher shot counts. Prompts are written in English and specify the model role, task, expected output or label set, any demonstrations, and the test input. In preliminary comparisons on dialect identification, machine-generated-text detection, toxicity detection, and machine translation, English prompts outperformed Arabic prompts. ChatGPT and the 7.1-billion-parameter instruction-tuned BLOOMZ are evaluated by free-form generation followed by simple post-processing. The supervised comparators are MARBERTv2 for NLU and AraT5 for NLG; both use learning rate 5e-5, with MARBERTv2 trained for 25 epochs and patience 5, and AraT5 for 10 epochs. Their best development-set checkpoints are used. Results are reported on the same 200 examples for direct comparison, and the paper additionally reports three-run averages on the full test sets for the supervised models. NLU is measured with macro-F1; NLG uses task-specific metrics.

  3. Knowl 3 — ChatGPT trails the Arabic NLU specialist on the overall benchmark

    empirical result

    On the ORCA aggregate score, ChatGPT rises from 37.98 in zero-shot to 51.28 at 10 shots; BLOOMZ peaks at 37.06 at 5 shots. MARBERTv2 scores 68.85 on the same 200-example evaluation sample and 71.00 on the full test sets. Thus, in this evaluation, the much smaller Arabic-specialized finetuned model substantially exceeds both prompted multilingual models overall. More shots do not reliably improve either prompted model. ChatGPT nevertheless has notable task-level exceptions: it reaches macro-F1 85.80 on paraphrase classification at 10 shots, exceeding MARBERTv2’s 65.98 on the sampled test set; scores 65.04 on gender classification, just above MARBERTv2’s 64.97; and achieves 53.49 on word-sense disambiguation at 3 shots, above MARBERTv2’s 36.31. These exceptions do not overturn the overall NLU pattern.

  4. Knowl 4 — NLG comparison across tasks

    data/table

    The table gives each prompted model’s best score among the reported shot settings and the AraT5 score on the shared 200-example test sample. Higher is better except for character error rate (CER), where lower is better. AraT5 leads on the non-code-switched tasks shown, while ChatGPT usually exceeds BLOOMZ; code-switched translation is the clear exception, where both prompted models exceed AraT5. The entries also show that BLOOMZ is notably stronger on QA and text rewriting, whereas ChatGPT is substantially stronger on translation and grammatical-error correction.

    Task Metric BLOOMZ best (shots) ChatGPT best (shots) AraT5 (200)
    Text rewriting BLEU 76.67 (0) 62.62 (10) 99.64
    Paraphrase generation BLEU 12.98 (0) 9.60 (10) 14.40
    Question generation BLEU 28.76 (0) 20.08 (5) 35.17
    Question answering SQuAD F1 76.04 (0) 54.14 (5) 81.45
    Summarization ROUGE-L 13.56 (0) 20.43 (5) 35.31
    News-title generation BLEU 1.20 (5) 4.72 (3) 7.72
    Diacritization CER, lower is better 0.51 (0) 0.05 (5) 0.03
    Transliteration CER, lower is better 0.42 (5 or 10) 0.23 (10) 0.18
    English to Arabic MT BLEU 12.54 (3) 23.74 (10) 27.12
    Spanish to Arabic MT BLEU 9.31 (5) 19.32 (10) 21.16
    French to Arabic MT BLEU 6.88 (0) 16.26 (10) 18.48
    Russian to Arabic MT BLEU 3.17 (5) 17.52 (3) 19.32
    Jordanian Arabic-English to English BLEU 11.56 (5) 40.88 (10) 5.56
    Algerian Arabic-French to French BLEU 28.61 (10) 37.95 (10) 17.49
    Grammatical-error correction M2Scorer F0.5 2.40 (0 or 5) 48.72 (0 or 3) 67.54
  5. Knowl 5 — Both GPT models perform better on MSA than dialectal Arabic

    empirical result

    For a dialect-focused NLU comparison, the authors select 11 ORCA tasks identified as involving more dialectal Arabic. An in-house MSA-versus-dialect classifier with approximately 88% F1 is used to select examples classified with at least 80% confidence, yielding averages of 119 dialectal and 78.82 Modern Standard Arabic (MSA) examples per task. Macro-F1 results show higher performance on MSA than on dialectal Arabic for ChatGPT on 8 of 11 tasks and for GPT-4 on 9 of 11. GPT-4 exceeds ChatGPT on 9 tasks for MSA and 7 for dialectal Arabic; the paper reports average GPT-4 improvements of 18.62% on MSA and 10.40% on dialectal tasks. In zero-shot country-level dialect identification, GPT-4 obtains 26.92 macro-F1, compared with ChatGPT’s zero-shot 13.86. These findings are evidence of a dialect-related performance gap in the selected evaluation data, not a causal account of how either model was trained.

  6. Knowl 6 — Dialect translation favors MSA, with GPT-4 gains on most varieties

    empirical result

    On machine translation from Arabic to English using MSA and Egyptian, Jordanian, Palestinian, Syrian, and Tunisian data from the Multi-dialectal Parallel Corpus, ChatGPT scores higher on MSA than on every dialect in all tested shot settings. Among the dialects, it performs best on Egyptian Arabic. In a separate zero-shot comparison using 50 examples per variety, GPT-4 outperforms ChatGPT on MSA and on all five dialects except Jordanian Arabic. Its largest advantages occur on Palestinian, Syrian, and Tunisian Arabic, which the authors characterize as relatively low-resource varieties. The paper suggests that differences in exposure to Arabic varieties may help explain these results, but presents this as a possible explanation rather than a measured cause.

  7. Knowl 7 — GPT-4 judgments align with human ratings on the evaluated generation tasks

    empirical result

    The human evaluation covers eight generation tasks: two code-switched translation tasks, three dialectal dialogue-generation tasks, machine translation, paraphrasing, and summarization. Six native Arabic speakers work in three annotator pairs; each pair rates 50 outputs per task on a four-level A–D scale based on how well the output fulfills the prompt. Inter-annotator agreement exceeds 0.6 for most tasks. Separately, GPT-4 rates outputs from BLOOMZ, ChatGPT, and GPT-4 on 25 randomly sampled examples from each of the eight datasets. GPT-4 generally rates its own outputs and ChatGPT’s above BLOOMZ’s, which is often rated D because it copies the input instead of following the instruction. At least one human evaluator agrees with GPT-4’s rating 71.5% of the time on average. This supports GPT-4 as a possible evaluator for the tasks studied, but does not establish that it can replace human assessment generally.

  8. Knowl 8 — ChatGPT translations are fluent but often fail to code-switch

    empirical result

    In a diagnostic human assessment, two annotators fluent in the relevant Arabic variety and English rate ChatGPT’s English-to-Arabic-English code-switched translations for MSA, Egyptian, and Moroccan Arabic. Ratings separately assess fluency, faithfulness to the source meaning, and actual code-switching, each on an A–D scale. Averaged across annotators and the three varieties, 43.3% of outputs receive A and 35.0% B for fluency; 91.7% receive A and 5.0% B for faithfulness. By contrast, only 10.0% receive A for code-switching, while 56.7% receive D. The annotators observe that ChatGPT often produces Arabic-script text rather than the requested mixed-language output, with this problem more prevalent for MSA than for Egyptian or Moroccan Arabic. The authors offer possible linguistic and lexical explanations, but do not establish them experimentally.

  9. Knowl 9 — ChatGPT frequently mislabels non-toxic Arabic as toxic

    empirical result

    Inspection of ChatGPT’s confusion matrices on Arabic abusive-language, hate-speech, and offensive-language classification shows a strong tendency to classify non-toxic examples as toxic, producing many false positives. The paper proposes two possible explanations: insufficiently diverse Arabic pretraining data, including limited coverage of some dialects, and safety supervision intended to reduce harmful outputs. These explanations are hypotheses; the reported finding is the false-positive tendency in the evaluated classification tasks.

  10. Knowl 10 — Evaluation scope limits how broadly the results can be generalized

    limitation

    Most model comparisons use a random sample of 200 test examples per dataset, so scores can differ from full-test-set results. ChatGPT’s behavior can also change as the service is updated; the reported results apply to the model snapshot evaluated in this study. GPT-4 generation and human or GPT-4 quality ratings cover selected tasks rather than the full benchmark. The study compares three multilingual LLMs—ChatGPT, GPT-4, and BLOOMZ—and two Arabic-specialized finetuned models, but omits other large or multilingual models for resource reasons. Open-domain dialectal dialogue generation is excluded from automatic NLG scoring because all models score close to zero BLEU and the metric is judged unsuitable for open-ended responses; those tasks instead receive human and GPT-4 evaluation.

Coverage note — The full NLU task-by-shot score matrix and the complete per-annotator human-rating counts are omitted because they are extensive breakdowns; the overall comparison, key task exceptions, and aggregate evaluator findings are retained.

References

  1. 1.Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan, and Kareem Darwish. 2020. Arabic Dialect Identification in the Wild. Proceedings of the Sixth Arabic Natural Language Processing Workshop.
  2. 2.Muhammad Abdul-Mageed, AbdelRahim Elmadany, and El Moatez Billah Nagoudi. 2021a. ARBERT & MARBERT: Deep bidirectional transformers for Arabic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7088–7105, Online. Association for Computational Linguistics.
  3. 3.Muhammad Abdul-Mageed, AbdelRahim Elmadany, and El Moatez Billah Nagoudi. 2021b. ARBERT & MARBERT: Deep bidirectional transformers for Arabic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7088–7105, Online. Association for Computational Linguistics.
  4. 4.Muhammad Abdul-Mageed, Chiyu Zhang, Houda Bouamor, and Nizar Habash. 2020a. NADI 2020: The first nuanced Arabic dialect identification shared task. In Proceedings of the Fifth Arabic Natural Language Processing Workshop, pages 97–110, Barcelona, Spain (Online). Association for Computational Linguistics.
  5. 5.Muhammad Abdul-Mageed, Chiyu Zhang, Azadeh Hashemi, and El Moatez Billah Nagoudi. 2020b. AraNet: A deep learning toolkit for Arabic social media. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pages 16–23, Marseille, France. European Language Resource Association.
  6. 6.Ibrahim Abu Farha and Walid Magdy. 2021. Benchmarking transformer-based language models for Arabic sentiment and sarcasm detection. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 21–31, Kyiv, Ukraine (Virtual). Association for Computational Linguistics.
  7. 7.Ali Saleh Alammary. 2022. Bert models for arabic text classification: A systematic review. Applied Sciences, 12(11).
  8. 8.Bashar Alhafni, Nizar Habash, and Houda Bouamor. 2022. User-centric gender rewriting. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 618–631, Seattle, United States. Association for Computational Linguistics.
  9. 9.Ali Alshehri, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2020. Understanding and detecting dangerous speech in social media. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pages 40–47, Marseille, France. European Language Resource Association.
  10. 10.Mohamed Seghir Hadj Ameur, Farid Meziane, and Ahmed Guessoum. 2019. Anetac: Arabic named entity transliteration and classification dataset. arXiv preprint arXiv:1907.03110.
  11. 11.Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. On the Cross-lingual Transferability of Monolingual Representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623–4637.
  12. 12.Ramy Baly, Mitra Mohtarami, James Glass, Lluís Màrquez, Alessandro Moschitti, and Preslav Nakov. 2018. Integrating stance detection and fact checking in a unified corpus. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 21–27.
  13. 13.Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung. 2023. A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity.
  14. 14.Houda Bouamor, Nizar Habash, and Kemal Oflazer. 2014. A multidialectal parallel corpus of arabic. In LREC, pages 1240–1245.
  15. 15.Houda Bouamor, Sabit Hassan, and Nizar Habash. 2019. The MADAR shared task on Arabic fine-grained dialect identification. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pages 199–207.
  16. 16.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  17. 17.Lingjiao Chen, Matei Zaharia, and James Zou. 2023a. How is chatgpt’s behavior changing over time? arXiv preprint arXiv:2307.09009.
  18. 18.Xuanting Chen, Junjie Ye, Can Zu, Nuo Xu, Rui Zheng, Minlong Peng, Jie Zhou, Tao Gui, Qi Zhang, and Xuanjing Huang. 2023b. How robust is gpt-3.5 to predecessors? a comprehensive study on language understanding tasks.
  19. 19.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  20. 20.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  21. 21.Hyung Won Chung, Le Hou, S. Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed Huai hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022a. Scaling instruction-finetuned language models. ArXiv, abs/2210.11416.
  22. 22.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022b. Scaling instruction-finetuned language models.
  23. 23.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. Xnli: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics.
  24. 24.Daniel Dahlmeier and Hwee Tou Ng. 2012. Better evaluation for grammatical error correction. In Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 568–572, Montréal, Canada. Association for Computational Linguistics.
  25. 25.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  26. 26.Mahmoud El-Haj. 2020. Habibi-a multi dialect multi national arabic song lyrics corpus. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 1318–1326.
  27. 27.Mohammed El-Razzaz, Mohamed Waleed Fakhr, and Fahima A Maghraby. 2021. Arabic gloss wsd using bert. Applied Sciences, 11(6):2567.
  28. 28.AbdelRahim Elmadany, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2022. Orca: A challenging benchmark for arabic language understanding. ArXiv, abs/2212.10758.
  29. 29.AbdelRahim Elmadany, El Moatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023. Orca: A challenging benchmark for arabic language understanding.
  30. 30.Kawin Ethayarajh. 2019. How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 55–65, Hong Kong, China. Association for Computational Linguistics.
  31. 31.Ali Fadel, Ibraheem Tuffaha, Bara’ Al-Jawarneh, and Mahmoud Al-Ayyoub. 2019. Arabic text diacritization using deep neural networks.
  32. 32.Ibrahim Abu Farha and Walid Magdy. 2020. From Arabic Sentiment Analysis to Sarcasm Detection: The ArSarcasm Dataset. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pages 32–39.
  33. 33.Yuan Gao, Ruili Wang, and Feng Hou. 2023. How to design translation prompts for chatgpt: An empirical study.
  34. 34.Bilal Ghanem, Jihen Karoui, Farah Benamara, Véronique Moriceau, and Paolo Rosso. 2019. IDAT@FIRE2019: Overview of the Track on Irony Detection in Arabic Tweets. . In Mehta P., Rosso P., Majumder P., Mitra M. (Eds.) Working Notes of the Forum for Information Retrieval Evaluation (FIRE 2019). CEUR Workshop Proceedings. In: CEUR-WS.org, Kolkata, India, December 12-15.
  35. 35.Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. Chatgpt outperforms crowd-workers for text-annotation tasks.
  36. 36.Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam, Kazi Samin, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. Xl-sum: Large-scale multilingual abstractive summarization for 44 languages.
  37. 37.Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are gpt models at machine translation? a comprehensive evaluation.
  38. 38.Haoyang Huang, Tianyi Tang, Dongdong Zhang, Wayne Xin Zhao, Ting Song, Yan Xia, and Furu Wei. 2023. Not all languages are created equal in llms: Improving multilingual capability by cross-lingual-thought prompting.
  39. 39.Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Xing Wang, and Zhaopeng Tu. 2023. Is chatgpt a good translator? yes with gpt-4 as the engine.
  40. 40.Jungo Kasai, Yuhei Kasai, Keisuke Sakaguchi, Yutaro Yamada, and Dragomir Radev. 2023. Evaluating gpt-4 and chatgpt on japanese medical licensing examinations.
  41. 41.Jude Khouja. 2020. Stance prediction and claim verification: An Arabic perspective. In Proceedings of the Third Workshop on Fact Extraction and VERification (FEVER), pages 8–17, Online. Association for Computational Linguistics.
  42. 42.Viet Dac Lai, Nghia Trung Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023. Chatgpt beyond english: Towards a comprehensive evaluation of large language models in multilingual learning.
  43. 43.Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq R. Joty, and J. Huang. 2023. A systematic study and comprehensive evaluation of chatgpt on benchmark datasets. ArXiv, abs/2305.18486.
  44. 44.Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. On the sentence embeddings from pre-trained language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9119–9130, Online. Association for Computational Linguistics.
  45. 45.Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 986–995.
  46. 46.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2022. Few-shot learning with multilingual language models.
  47. 47.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint arXiv:1907.11692.
  48. 48.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Computational Linguistics.
  49. 49.Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018. SemEval-2018 task 1: Affect in tweets. In Proceedings of The 12th International Workshop on Semantic Evaluation, pages 1–17, New Orleans, Louisiana. Association for Computational Linguistics.
  50. 50.Behrang Mohit, Alla Rozovskaya, Nizar Habash, Wajdi Zaghouani, and Ossama Obeid. 2014. The first QALB shared task on automatic text correction for Arabic. In Proceedings of the EMNLP 2014 Workshop on Arabic Natural Language Processing (ANLP), pages 39–47, Doha, Qatar. Association for Computational Linguistics.
  51. 51.Hamdy Mubarak, Kareem Darwish, Walid Magdy, Tamer Elsayed, and Hend Al-Khalifa. 2020. Overview of OSACT4 Arabic offensive language detection shared task. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pages 48–52, Marseille, France. European Language Resource Association.
  52. 52.Hamdy Mubarak, Sabit Hassan, and Ahmed Abdelali. 2021. Adult content detection on Arabic Twitter: Analysis and experiments. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 136–144, Kyiv, Ukraine (Virtual). Association for Computational Linguistics.
  53. 53.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2022a. Crosslingual generalization through multitask finetuning.
  54. 54.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Rose Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir R. Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2022b. Crosslingual generalization through multitask finetuning. ArXiv, abs/2211.01786.
  55. 55.El Moatez Billah Nagoudi, Muhammad Abdul-Mageed, AbdelRahim Elmadany, Alcides Alcoba Inciarte, and Md Tawkat Islam Khondaker. 2022a. Jasmine: Arabic gpt models for few-shot learning. arXiv preprint arXiv:2212.10755.
  56. 56.El Moatez Billah Nagoudi, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. 2022b. AraT5: Text-to-text transformers for Arabic language generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 628–647, Dublin, Ireland. Association for Computational Linguistics.
  57. 57.El Moatez Billah Nagoudi, AbdelRahim Elmadany, Muhammad Abdul-Mageed, and Tariq Alhindi. 2020. Machine generation and detection of Arabic manipulated and fake news. In Proceedings of the Fifth Arabic Natural Language Processing Workshop, pages 69–84, Barcelona, Spain (Online). Association for Computational Linguistics.
  58. 58.Tarek Naous, Zahraa Bassyouni, Bassel Mousi, Hazem Hajj, Wassim El Hajj, and Khaled Shaban. 2023. Open-domain response generation in low-resource settings using self-supervised pre-training of warm-started transformers. ACM Transactions on Asian and Low-Resource Language Information Processing, 22(4):1–12.
  59. 59.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial nli: A new benchmark for natural language understanding.
  60. 60.Reham Omar, Omij Mangukiya, Panos Kalnis, and Essam Mansour. 2023. Chatgpt versus traditional question answering for knowledge graphs: Current status and future directions towards knowledge graph chatbots.
  61. 61.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  62. 62.Keqin Peng, Liang Ding, Qihuang Zhong, Li Shen, Xuebo Liu, Min Zhang, Yuanxin Ouyang, and Dacheng Tao. 2023. Towards making the most of chatgpt for machine translation.
  63. 63.Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023a. Is chatgpt a general-purpose natural language processing task solver?
  64. 64.Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023b. Is chatgpt a general-purpose natural language processing task solver? CoRR, abs/2302.06476.
  65. 65.Michael V. Reiss. 2023. Testing the reliability of chatgpt for text annotation and classification: A cautionary remark.
  66. 66.Yves Scherrer. 2020. TaPaCo: A corpus of sentential paraphrases for 73 languages. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 6868–6873, Marseille, France. European Language Resources Association.
  67. 67.Haitham Seelawi, Ahmad Mustafa, Hesham Al-Bataineh, Wael Farhan, and Hussein T Al-Natsheh. 2019. Nsurl-2019 task 8: Semantic question similarity in arabic. In Proceedings of The First International Workshop on NLP Solutions for Under Resourced Languages (NSURL 2019) co-located with ICNLSP 2019-Short Papers, pages 1–8.
  68. 68.Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023. In chatgpt we trust? measuring and characterizing the reliability of chatgpt.
  69. 69.Yiming Tan, Dehai Min, Yu Li, Wenbo Li, Nan Hu, Yongrui Chen, and Guilin Qi. 2023. Evaluation of chatgpt as a question answering system for answering complex questions.
  70. 70.NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022. No language left behind: Scaling human-centered machine translation.
  71. 71.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. Glue: A multi-task benchmark and analysis platform for natural language understanding.
  72. 72.Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. 2022a. Adversarial glue: A multi-task benchmark for robustness evaluation of language models.
  73. 73.Jindong Wang, Xixu Hu, Wenxin Hou, Hao Chen, Runkai Zheng, Yidong Wang, Linyi Yang, Haojun Huang, Wei Ye, Xiubo Geng, Binxin Jiao, Yue Zhang, and Xing Xie. 2023a. On the robustness of chatgpt: An adversarial and out-of-distribution perspective.
  74. 74.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022b. Self-instruct: Aligning language model with self generated instructions. ArXiv, abs/2212.10560.
  75. 75.Zengzhi Wang, Qiming Xie, Zixiang Ding, Yi Feng, and Rui Xia. 2023b. Is chatgpt a good sentiment analyzer? a preliminary study.
  76. 76.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2021. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652.
  77. 77.Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C. Schmidt. 2023. A prompt pattern catalog to enhance prompt engineering with chatgpt. ArXiv, abs/2302.11382.
  78. 78.Jeff Wu, Long Ouyang, Daniel M. Ziegler, Nisan Stiennon, Ryan Lowe, Jan Leike, and Paul Christiano. 2021. Recursively summarizing books with human feedback.
  79. 79.Jiageng Wu, Xian Wu, Zhaopeng Qiu, Minghui Li, Yefeng Zheng, and Jie Yang. 2023a. Qualifying chinese medical licensing examination with knowledge enhanced generative pre-training model.
  80. 80.Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, and Alham Fikri Aji. 2023b. Lamini-lm: A diverse herd of distilled models from large-scale instructions.
  81. 81.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 483–498, Online. Association for Computational Linguistics.
  82. 82.Omar F Zaidan and Chris Callison-Burch. 2014. Arabic Dialect Identification . Computational Linguistics, 40(1):171–202.
  83. 83.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  84. 84.Shen Zheng, Jie Huang, and Kevin Chen-Chuan Chang. 2023. Why does chatgpt fall short in answering questions faithfully?
  85. 85.Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2023a. Can chatgpt understand too? a comparative study on chatgpt and fine-tuned bert. ArXiv, abs/2302.10198.
  86. 86.Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2023b. Can chatgpt understand too? a comparative study on chatgpt and fine-tuned bert.
  87. 87.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023. Lima: Less is more for alignment.
  88. 88.Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2023. Multilingual machine translation with large language models: Empirical results and analysis.
  89. 89.Caleb Ziems, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2023. Can large language models transform computational social science? CoRR, abs/2305.03514.
  90. 90.Michal Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016. The united nations parallel corpus v1. 0. In Lrec.

Citation

MLA
Khondaker, M. T. I., et al. “GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 220–47, https://doi.org/10.18653/v1/2023.emnlp-main.16.
APA
Khondaker, M. T. I., Waheed, A., Nagoudi, E.-M.-B., & Abdul-Mageed, M. (2023). GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 220–247. https://doi.org/10.18653/v1/2023.emnlp-main.16
Chicago
Khondaker, M. T. I., A. Waheed, E.-M.-B. Nagoudi, and M. Abdul-Mageed. 2023. “GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 220–47. https://doi.org/10.18653/v1/2023.emnlp-main.16.
Harvard
Khondaker, M.T.I. et al. (2023) “GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 220–247. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.16.
Vancouver
1. Khondaker MTI, Waheed A, Nagoudi E-M-B, Abdul-Mageed M (2023) GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 220–247

BibTeX

@inproceedings{khondaker-etal-2023-gptaraeval,
    title = "{GPTA}ra{E}val: A Comprehensive Evaluation of {C}hat{GPT} on {A}rabic {NLP}",
    author = "Khondaker, Md Tawkat Islam  and
      Waheed, Abdul  and
      Nagoudi, El Moatez Billah  and
      Abdul-Mageed, Muhammad",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.16/",
    doi = "10.18653/v1/2023.emnlp-main.16",
    pages = "220--247"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/