GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP
Md. Tawkat Islam KhondakerAbdul WaheedEl Moatez Billah NagoudiMuhammad Abdul-Mageed
Demonstrates through a large-scale evaluation across 44 tasks and over 60 datasets that ChatGPT and GPT-4 consistently lag behind smaller, fine-tuned dedicated models on Arabic natural language processing, particularly when handling regional dialectal varieties compared to Modern Standard Arabic.
Large language models such as ChatGPT have demonstrated strong capabilities across standard English-language benchmarks, driving rapid adoption across global industries. However, their effectiveness in linguistically rich and complex languages like Arabic—spoken by more than 450 million people across diverse regional dialects—remains largely unverified at scale. The article systematically evaluates ChatGPT across a wide range of Arabic natural language understanding and generation tasks to determine whether general-purpose commercial models can reliably serve Arabic-speaking users without task-specific customization.
The article conducts an extensive empirical assessment across 44 distinct Arabic language understanding and text generation tasks spanning more than 60 datasets. The authors test ChatGPT against an instruction-tuned multilingual baseline (BLOOMZ) and two dedicated, smaller Arabic models that were specifically finetuned for Arabic understanding and generation (MARBERTV2 and AraT5). Evaluations were run across zero-shot and few-shot prompt setups using English instruction templates, complemented by targeted evaluations of country-level dialects and qualitative assessments conducted by both native human annotators and GPT-4.
The evaluation reveals several critical findings. First, despite its massive scale and general multilingual pretraining, ChatGPT is consistently outperformed by significantly smaller, Arabic-dedicated finetuned models across almost all understanding and generation benchmarks. On overall language understanding, the dedicated finetuned model achieves an aggregate macro-F1 score of approximately 69% to 71%, whereas ChatGPT peaks at roughly 51% even with ten-shot prompting. Second, both ChatGPT and GPT-4 demonstrate a substantial performance drop when processing Dialectal Arabic compared to Modern Standard Arabic, with GPT-4 outperforming ChatGPT by an average of about 19% on standard Arabic and 10% on regional dialects. Third, in text generation tasks such as machine translation, summarization, and grammatical error correction, task-dedicated models consistently achieve higher accuracy and lower error rates, although ChatGPT outperforms other models on code-switched inputs involving mixtures of Arabic and English or French. Finally, human and automated quality assessments reveal that GPT-4 ratings agree with human evaluators approximately 71.5% of the time.
These findings indicate that relying on general-purpose, out-of-the-box language models for Arabic applications introduces notable operational, quality, and compliance risks. Deploying standard commercial models in production without adaptation could lead to degraded user experiences, particularly for populations communicating in regional dialects rather than formal modern Arabic. Additionally, the analysis revealed that ChatGPT frequently exhibits high false-positive rates on content moderation tasks, often incorrectly classifying benign dialectal text as toxic. This tendency poses a risk of unnecessary content censorship or flawed moderation workflows if deployed without safeguards.
Organizations seeking to implement Arabic artificial intelligence solutions should prioritize specialized, finetuned models or explore hybrid architectures rather than relying entirely on generic foundation models. When commercial models like GPT-4 are utilized, practitioners should establish robust evaluation pipelines and restrict deployment to Modern Standard Arabic workflows until dialect handling improves. The findings are based on controlled test samples and fixed model snapshots from early 2023, meaning performance may evolve as commercial models update; nevertheless, the results provide high confidence that dedicated domain adaptation remains essential for reliable Arabic natural language processing.
- Paper: AraBERT: Transformer-based Model for Arabic Language Understanding, Wissam Antoun et al. (2020). AraBERT establishes how Arabic-specific pretraining and linguistic tailoring can outperform generic multilingual models, providing useful context for GPTAraEval’s comparisons with dedicated Arabic systems.
- Paper: AceGPT, Localizing Large Language Models in Arabic, Huang Huang et al. (2024). AceGPT follows the evaluation’s case for Arabic adaptation by building and testing localized Arabic models, including culturally aligned training and evaluation.
