AraBERT: Transformer-based Model for Arabic Language Understanding

Wissam AntounFady BalyHazem M. Hajj

article2020OSACT1,507 citations

Introduces AraBERT, an Arabic-specific transformer model pretrained on large-scale text that outperforms multilingual BERT across sentiment analysis, named entity recognition, and question answering benchmarks.

Listen

Natural language processing in Arabic has historically lagged behind English due to the language's complex grammatical structures, rich morphology, and a shortage of large, standardized datasets. While generic multilingual models cover Arabic alongside dozens of other languages, their limited vocabulary and insufficient Arabic training data often lead to suboptimal performance in real-world business and technological applications.

The article demonstrates the development and evaluation of AraBERT, a specialized transformer-based language model trained exclusively on Arabic text. The primary objective was to establish a dedicated, high-performing foundation model capable of advancing state-of-the-art results across several core Arabic natural language understanding tasks.

To build the model, the authors trained a 110-million parameter architecture on a newly compiled 24-gigabyte dataset comprising 70 million sentences extracted from news sources across multiple Arab countries. To address Arabic's specific structure, words were pre-segmented into prefixes, stems, and suffixes prior to building a tailored 64,000-token vocabulary. The system was then evaluated across three core tasks using eight benchmark datasets: sentiment analysis across multiple dialects, named entity recognition, and question answering.

The findings show that AraBERT outperformed both Google's multilingual BERT and previous specialized baselines across almost all benchmarks. In sentiment analysis, it achieved substantial accuracy gains across both Modern Standard Arabic and regional dialects, including an improvement from 52.4% to 59.4% on Levantine tweets. On named entity recognition, AraBERT achieved a state-of-the-art F1 score of 84.2%, outperforming the previous baseline of 81.7%. For question answering, it achieved a 93.0% sentence match rate, improving over the previous 90.0% benchmark, although exact-match scores were impacted by minor variations in prepositions and introductory phrasing.

These results indicate that dedicated, language-specific models provide superior comprehension and operational accuracy compared to general multilingual systems. Additionally, AraBERT requires approximately 300 megabytes less storage than multilingual BERT, offering improved performance at lower computational and hosting overhead. This makes it an effective foundation for enterprise text processing, search, customer feedback analysis, and automated information retrieval in Arabic.

Organizations developing Arabic-language applications should adopt AraBERT as a baseline foundation model rather than relying on general multilingual alternatives. However, developers should note that pre-segmenting text benefits sentiment and question-answering tasks, whereas the non-segmented version performs better for entity recognition. Future research and development should focus on removing external tokenizer dependencies and expanding pre-training to better cover colloquial regional dialects.

Cover for AraBERT: Transformer-based Model for Arabic Language Understanding

Abstract

The Arabic language is a morphologically rich language with relatively few resources and a less explored syntax compared to English. Given these limitations, Arabic Natural Language Processing (NLP) tasks like Sentiment Analysis (SA), Named Entity Recognition (NER), and Question Answering (QA), have proven to be very challenging to tackle. Recently, with the surge of transformers based models, language-specific BERT based models have proven to be very efficient at language understanding, provided they are pre-trained on a very large corpus. Such models were able to set new standards and achieve state-of-the-art results for most NLP tasks. In this paper, we pre-trained BERT specifically for the Arabic language in the pursuit of achieving the same success that BERT did for the English language. The performance of AraBERT is compared to multilingual BERT from Google and other state-of-the-art approaches. The results showed that the newly developed AraBERT achieved state-of-the-art performance on most tested Arabic NLP tasks. The pretrained araBERT models are publicly available on this https URL hoping to encourage research and applications for Arabic NLP.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Evolution of Word Embeddings
  • 2.2. Non-contextual Representations for Arabic
  • 2.3. Contextualized Representations for Arabic
  • 3. ARABERT: Methodology
  • 3.1. Pre-training Setup
  • 3.2. Pre-training Dataset
  • 3.3. Sub-Word Units Segmentation
  • 3.4. Fine-tuning
  • 4. Evaluation
  • 4.1. Sentiment Analysis
  • 4.2. Named Entity Recognition
  • 4.3. Question Answering
  • 5. Experiments
  • 5.1. Experimental Setup
  • 5.2. Results
  • 5.3. Discussion
  • 6. Conclusion
  • 7. Acknowledgments
  • 8. References

Knowls

  1. Knowl 1 — AraBERT Model Architecture and Pre-training Objectives

    model/method

    AraBERT is an Arabic-specific language representation model based on the stacked Bidirectional Transformer Encoder architecture using the standard BERT-base structural configuration:

    • Transformer Encoder Layers: 12 blocks
    • Hidden Representation Size: 768 dimensions
    • Self-Attention Heads: 12 heads
    • Maximum Sequence Length: 512 tokens
    • Parameter Count: Approximately 110 million parameters

    AraBERT is pre-trained using two simultaneous self-supervised objectives:

    1. Masked Language Modeling (MLM) with Whole-Word Masking: A subset representing 15% of all input tokens is selected for replacement. Whole-word masking is enforced such that whenever any sub-word token of a word is selected, all constituent sub-word tokens forming that entire word are masked simultaneously. Of the selected tokens, 80% are replaced with the special [MASK] token, 10% are replaced with a random vocabulary token, and 10% remain unmodified.
    2. Next Sentence Prediction (NSP): A binary classification task where the model receives sentence pairs (A,B)(A, B) and predicts whether sentence BB is the actual contiguous next sentence following sentence AA in the source text or a randomly sampled sentence from the corpus.
  2. Knowl 2 — Morphological Pre-segmentation and Sub-word Tokenization Pipeline

    model/method

    Arabic exhibits high lexical sparsity due to its complex concatenative morphology, where non-intrinsic affixes—such as the definite article "الـ" (AlAl-), conjunctions, and possessive pronominal suffixes—attach directly to base stems without spaces. Direct application of sub-word tokenization causes duplicate vocabulary entries for words with and without these affixes.

    To address this, the AraBERT framework introduces a morphology-aware tokenization pipeline alongside an unsegmented variant:

    • AraBERT (v1): Input text is first morphologically segmented using the Farasa segmenter into constituent prefixes, stems, and suffixes (for example, "اللغة - Alloga" is segmented into "Al+ log +a"). An unsupervised SentencePiece tokenizer is then trained in unigram mode over the segmented pre-training text to construct a sub-word vocabulary of approximately 60,000 tokens.
    • AraBERTv0.1: SentencePiece is trained directly on unsegmented text in unigram mode, removing the dependency on external morphological segmenters during tokenization.
    • Vocabulary Allocation: Both versions utilize a total vocabulary size of 64,000 tokens, which includes approximately 4,000 reserved unused tokens to facilitate subsequent domain-specific pre-training.
  3. Knowl 3 — AraBERT Pre-training Corpus Construction

    experimental setup

    The pre-training dataset for AraBERT was compiled from three large Arabic text sources to achieve regional and topical diversity across Modern Standard Arabic (MSA):

    1. Manual Web Scrapes: Scraped articles from multiple Arabic news websites.
    2. 1.5 Billion Words Arabic Corpus: A contemporary corpus comprising over 5 million news articles spanning ten major news organizations across 8 Arab countries.
    3. OSIAN (Open Source International Arabic News Corpus): A corpus containing approximately 3.5 million articles (~1 billion tokens) collected from 31 news sources across 24 Arab countries.

    Following sentence-level deduplication, the finalized pre-training corpus contains 70 million sentences, totaling approximately 24 GB of text. Words containing Latin characters (such as scientific terms, technical vocabulary, and foreign named entities) were intentionally preserved in their original form to avoid semantic information loss.

  4. Knowl 4 — AraBERT Pre-training Optimization and Two-Phase Sequence Schedule

    experimental setup

    AraBERT was pre-trained using TensorFlow on a Google Cloud TPUv2-8 pod for 1,250,000 total optimization steps over 4 days (amounting to 27 epochs over the 70 million sentence corpus):

    • Two-Phase Sequence Length Schedule: To accelerate pre-training throughput, the first 900,000 steps were trained with a maximum sequence length of 128 tokens and a batch size of 512. The remaining 350,000 steps were trained with a maximum sequence length of 512 tokens and a batch size of 128.
    • Optimizer and Hyperparameters: Optimized using the Adam optimizer with a learning rate of 1imes10−41 imes 10^{-4}.
    • Data Duplication: Pre-training input data was converted to sharded TFRecords with a duplication factor of 10, a random seed of 34, and a masking probability of 15%.
    • Stopping Criterion: Pre-training termination was determined by tracking the convergence and validation performance on downstream Arabic NLU evaluation tasks.
  5. Knowl 5 — Downstream Fine-Tuning Architectures for Arabic NLU

    model/method

    AraBERT is adapted to downstream Arabic Natural Language Understanding tasks via task-specific output layers that are trained jointly with the pre-trained transformer weights:

    • Sequence Classification (Sentiment Analysis): The contextual hidden representation vector h[CLS]∈R768\mathbf{h}_{\text{[CLS]}} \in \mathbb{R}^{768} of the prepended [CLS] token is passed through a single feed-forward classification layer followed by a Softmax activation to generate the class probability distribution. All layers are fine-tuned jointly by minimizing cross-entropy loss.
    • Named Entity Recognition (NER): Framed as token-level sequence classification using the IOB2 tagging scheme. When words are partitioned into multiple sub-tokens by the SentencePiece tokenizer, only the contextual embedding of the first sub-token of each word is passed into the linear entity classifier to predict the entity tag (B-, I-, or O).
    • Question Answering (Span Extraction): Given a concatenated sequence of question tokens and passage tokens separated by [SEP], the final output embedding hi\mathbf{h}_i of each passage token ii is evaluated by two separate linear projection vectors, wstart\mathbf{w}_{\text{start}} and wend\mathbf{w}_{\text{end}}. The probability that token ii is the start or end of the answer span is computed via Softmax over all passage tokens: Pstart(i)=exp⁡(wstartThi)∑jexp⁡(wstartThj),Pend(i)=exp⁡(wendThi)∑jexp⁡(wendThj)P_{\text{start}}(i) = \frac{\exp(\mathbf{w}_{\text{start}}^T \mathbf{h}_i)}{\sum_j \exp(\mathbf{w}_{\text{start}}^T \mathbf{h}_j)}, \quad P_{\text{end}}(i) = \frac{\exp(\mathbf{w}_{\text{end}}^T \mathbf{h}_i)}{\sum_j \exp(\mathbf{w}_{\text{end}}^T \mathbf{h}_j)} The predicted answer is selected by maximizing Pstart(i)⋅Pend(j)P_{\text{start}}(i) \cdot P_{\text{end}}(j) subject to the constraint j≥ij \ge i.
  6. Knowl 6 — Benchmark Performance of AraBERT on Arabic NLU Downstream Tasks

    data/table

    AraBERT was evaluated across three core Arabic NLU task domains—Sentiment Analysis (SA), Named Entity Recognition (NER), and Question Answering (QA)—and compared against multilingual BERT (mBERT) and prior state-of-the-art (SOTA) task-specific architectures:

    Task Metric Prev. SOTA mBERT AraBERTv0.1 AraBERTv1
    SA (HARD) Accuracy (%) 95.7 95.7 96.2 96.1
    SA (ASTD) Accuracy (%) 86.5 80.1 92.2 92.6
    SA (ArSenTD-Lev) Accuracy (%) 52.4 51.0 58.9 59.4
    SA (AJGT) Accuracy (%) 92.6 83.6 93.1 93.8
    SA (LABR) Accuracy (%) 87.5 83.0 85.9 86.7
    NER (ANERcorp) Macro-F1 81.7 78.4 84.2 81.9
    QA (ARCD) Exact Match (%) – 34.2 30.1 30.6
    QA (ARCD) Macro-F1 – 61.3 61.2 62.7
    QA (ARCD) Sentence Match (%) – 90.0 93.0 92.0

    Previous SOTA baselines: hULMonA for HARD, ASTD, and ArSenTD-Lev; multi-channel CNN models for AJGT and LABR; BiLSTM-CRF for ANERcorp; and mBERT for ARCD.

    AraBERT establishes a new state-of-the-art across nearly all evaluated benchmarks, significantly outperforming mBERT due to its larger single-language pre-training corpus (24 GB vs. 4.3 GB Wikipedia text in mBERT) and dedicated Arabic vocabulary (64k tokens vs. 2k Arabic sub-words in mBERT).

  7. Knowl 7 — Generalization of AraBERT to Dialectal Arabic Sentiment Analysis

    empirical result

    Despite being pre-trained exclusively on Modern Standard Arabic (MSA) text corpora, AraBERT generalizes effectively when fine-tuned on diverse Dialectal Arabic (DA) sentiment datasets:

    • Levantine Dialect (ArSenTD-Lev): AraBERTv1 achieves 59.4% accuracy (and AraBERTv0.1 achieves 58.9%), surpassing the previous state-of-the-art hULMonA (52.4%) and mBERT (51.0%) by up to 7.0 absolute percentage points.
    • Jordanian Dialect (AJGT): AraBERTv1 attains 93.8% accuracy, outperforming the prior CNN SOTA (92.6%) and substantially outperforming mBERT (83.6%).
    • Egyptian and MSA Dialects (ASTD-B): AraBERTv1 reaches 92.6% accuracy, improving by 6.1 percentage points over hULMonA (86.5%) and 12.5 percentage points over mBERT (80.1%).
    • Mixed MSA and Dialectal Reviews (HARD): AraBERT models achieve 96.1%–96.2% accuracy, exceeding both hULMonA and mBERT (95.7%).

    These results indicate that representations learned from large-scale MSA pre-training provide strong semantic transfer capabilities to unseen dialectal variations of Arabic.

  8. Knowl 8 — Interaction of Morphological Segmentation with Arabic Named Entity Recognition

    empirical result

    On the ANERcorp Named Entity Recognition benchmark, the unsegmented model variant AraBERTv0.1 achieved a macro-F1 score of 84.2, setting a new state of the art by outperforming the prior BiLSTM-CRF baseline (81.7 macro-F1) and multilingual BERT (78.4 macro-F1).

    However, the morphologically pre-segmented variant AraBERTv1 achieved a lower macro-F1 score of 81.9 (comparable to the BiLSTM-CRF baseline). This performance difference arises from an alignment conflict between morphological prefix stripping and IOB2 entity tagging:

    • When a prefixed entity noun such as "الجامعة" (labeled B-ORG) is pre-segmented into prefix and stem ("الـ + جامعة"), the entity start tag (B-ORG) is assigned to the detached prefix "الـ" (AlAl-), while the semantic noun "جامعة" receives the continuation tag I-ORG.
    • Because prefixes like "الـ" are ubiquitous across non-entity nouns throughout the language, designating generic detached prefixes as the starting cue (B- label) injects ambiguity and degrades the sequence tagger's ability to identify entity boundaries.
  9. Knowl 9 — Span Boundary and Sentence Retrieval Dynamics in Arabic Question Answering

    empirical result

    On the Arabic Reading Comprehension Dataset (ARCD), AraBERTv1 achieved a macro-F1 score of 62.7 (compared to 61.3 for multilingual BERT) and improved Sentence Match (SM)—the percentage of predictions located within the correct ground-truth sentence—by 2.0 to 3.0 percentage points absolute (92.0% for AraBERTv1 and 93.0% for AraBERTv0.1 versus 90.0% for mBERT).

    Despite higher F1 and Sentence Match scores, Exact Match (EM) scores for AraBERT (30.1% for v0.1 and 30.6% for v1) were lower than mBERT (34.2%). Error analysis revealed two primary sources for the EM reduction:

    1. Preposition Shifts: The predicted answer span differed from the gold reference by single leading prepositions without affecting semantic correctness (for example, predicting "سان فرانسيسكو - San Francisco" instead of the ground truth "في سان فرانسيسكو - In San Francisco").
    2. Introductory / Copula Word Omission: Predictions frequently omitted leading introductory context while capturing the complete factual answer (for example, predicting "جمهورية فيدرالية - A federal republic" instead of "النمسا هي جمهورية فيدرالية - Austria is a federal republic").

Coverage note — No substantial contributed material was omitted from the extracted knowls.

References

  1. 1.Abdelali, A., Darwish, K., Durrani, N., and Mubarak, H. (2016). Farasa: A fast and furious segmenter for arabic. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations, pages 11–16.
  2. 2.Abdul-Mageed, M., Alhuzali, H., and Elaraby, M. (2018). You tweet what you speak: A city-level dataset of arabic dialects. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018).
  3. 3.Abu Farha, I. and Magdy, W. (2019). Mazajak: An online Arabic sentiment analyser. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pages 192–198, Florence, Italy, August. Association for Computational Linguistics.
  4. 4.Adiwardana, D., Luong, M.-T., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., and Le, Q. V. (2020). Towards a human-like open-domain chatbot.
  5. 5.Al Sallab, A., Hajj, H., Badaro, G., Baly, R., El-Hajj, W., and Shaban, K. (2015). Deep learning models for sentiment analysis in arabic. In Proceedings of the second workshop on Arabic natural language processing, pages 9–17.
  6. 6.Al-Sallab, A., Baly, R., Hajj, H., Shaban, K. B., El-Hajj, W., and Badaro, G. (2017). Aroma: A recursive deep learning model for opinion mining in arabic as a low resource language. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 16(4):1–20.
  7. 7.Alomari, K. M., ElSherif, H. M., and Shaalan, K. (2017). Arabic tweets sentimental analysis using machine learning. In International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems, pages 602–610. Springer.
  8. 8.Aly, M. and Atiya, A. (2013). LABR: A large scale Arabic book reviews dataset. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 494–498, Sofia, Bulgaria, August. Association for Computational Linguistics.
  9. 9.Badaro, G., Baly, R., Hajj, H., Habash, N., and El-Hajj, W. (2014). A large scale arabic sentiment lexicon for arabic opinion mining. In Proceedings of the EMNLP 2014 workshop on arabic natural language processing (ANLP), pages 165–173.
  10. 10.Baly, R., Hajj, H., Habash, N., Shaban, K. B., and El-Hajj, W. (2017). A sentiment treebank and morphologically enriched recursive deep models for effective sentiment analysis in arabic. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 16(4):1–21.
  11. 11.Baly, R., Khaddaj, A., Hajj, H., El-Hajj, W., and Shaban, K. B. (2018). Arsentd-lev: A multi-topic corpus for target-based sentiment analysis in arabic levantine tweets. In OSACT 3: The 3rd Workshop on Open-Source Arabic Corpora and Processing Tools, page 37.
  12. 12.Benajiba, Y. and Rosso, P. (2007). Anersys 2.0: Conquering the ner task for the arabic language by combining the maximum entropy with pos-tag information. In IICAI, pages 1814–1823.
  13. 13.Bojanowski, P., Grave, E., Joulin, A., and Mikolov, T. (2017). Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146.
  14. 14.Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V. (2019). Unsupervised cross-lingual representation learning at scale.
  15. 15.Dahou, A., Elaziz, M. A., Zhou, J., and Xiong, S. (2019a). Arabic sentiment classification using convolutional neural network and differential evolution algorithm. Computational intelligence and neuroscience, 2019.
  16. 16.Dahou, A., Xiong, S., Zhou, J., and Elaziz, M. A. (2019b). Multi-channel embedding convolutional neural network model for arabic sentiment classification. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), 18(4):1–23.
  17. 17.de Vries, W., van Cranenburgh, A., Bisazza, A., Caselli, T., van Noord, G., and Nissim, M. (2019). Bertje: A dutch bert model. arXiv preprint arXiv:1912.09582.
  18. 18.DeepsetAI. (2019). Open sourcing german bert.
  19. 19.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
  20. 20.El Bazi, I. and Laachfoubi, N. (2019). Arabic named entity recognition using deep learning approach. International Journal of Electrical & Computer Engineering (2088-8708), 9(3).
  21. 21.El-Khair, I. A. (2016). 1.5 billion words arabic corpus. arXiv preprint arXiv:1611.04033.
  22. 22.ElJundi, O., Antoun, W., El Droubi, N., Hajj, H., El-Hajj, W., and Shaban, K. (2019). hulmona: The universal language model in arabic. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pages 68–77.
  23. 23.Elnagar, A., Khalifa, Y. S., and Einea, A. (2018). Hotel arabic-reviews dataset construction for sentiment analysis applications. In Intelligent Natural Language Processing: Trends and Applications, pages 35–52. Springer.
  24. 24.Erdmann, A., Zalmout, N., and Habash, N. (2018). Addressing noise in multidialectal word embeddings. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 558–565.
  25. 25.Howard, J. and Ruder, S. (2018). Universal language model fine-tuning for text classification. arXiv preprint arXiv:1801.06146.
  26. 26.Huang, Z., Xu, W., and Yu, K. (2015). Bidirectional lstm-crf models for sequence tagging. arXiv preprint arXiv:1508.01991.
  27. 27.Kudo, T. (2018). Subword regularization: Improving neural network translation models with multiple subword candidates.
  28. 28.Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Kelcey, M., Devlin, J., Lee, K., Toutanova, K. N., Jones, L., Chang, M.-W., Dai, A., Uszkoreit, J., Le, Q., and Petrov, S. (2019). Natural questions: a benchmark for question answering research. Transactions of the Association of Computational Linguistics.
  29. 29.Lafferty, J. D., McCallum, A., and Pereira, F. C. (2001). Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In ICML.
  30. 30.Lample, G., Ballesteros, M., Subramanian, S., Kawakami, K., and Dyer, C. (2016). Neural architectures for named entity recognition. arXiv preprint arXiv:1603.01360.
  31. 31.Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2019). Albert: A lite bert for selfsupervised learning of language representations.
  32. 32.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach.
  33. 33.Martin, L., Muller, B., Suárez, P. J. O., Dupont, Y., Romary, L., Éric Villemonte de la Clergerie, Seddah, D., and Sagot, B. (2019). Camembert: a tasty french language model.
  34. 34.Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. (2013). Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pages 3111–3119.
  35. 35.Mikolov, T., Grave, E., Bojanowski, P., Puhrsch, C., and Joulin, A. (2017). Advances in pre-training distributed word representations. arXiv preprint arXiv:1712.09405.
  36. 36.Mozannar, H., Maamary, E., El Hajal, K., and Hajj, H. (2019). Neural arabic question answering. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pages 108–118.
  37. 37.Nabil, M., Aly, M., and Atiya, A. (2015). ASTD: Arabic sentiment tweets dataset. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 2515–2519, Lisbon, Portugal, September. Association for Computational Linguistics.
  38. 38.Pennington, J., Socher, R., and Manning, C. (2014). Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), pages 1532–1543.
  39. 39.Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. (2018). Deep contextualized word representations. In Proceedings of NAACLHLT, pages 2227–2237.
  40. 40.Polignano, M., Basile, P., de Gemmis, M., Semeraro, G., and Basile, V. (2019). AlBERTo: Italian BERT Language Understanding Model for NLP Challenging Tasks Based on Tweets. In Proceedings of the Sixth Italian Conference on Computational Linguistics (CLiC-it 2019), volume 2481. CEUR.
  41. 41.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2019). Exploring the limits of transfer learning with a unified text-to-text transformer.
  42. 42.Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016). Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250.
  43. 43.Ratnaparkhi, A. (1998). Maximum entropy models for natural language ambiguity resolution.
  44. 44.Sang, E. F. and De Meulder, F. (2003). Introduction to the conll-2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050.
  45. 45.Soliman, A. B., Eissa, K., and El-Beltagy, S. R. (2017). Aravec: A set of arabic word embedding models for use in arabic nlp. Procedia Computer Science, 117:256–265.
  46. 46.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need.
  47. 47.Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q. V. (2019). Xlnet: Generalized autoregressive pretraining for language understanding.
  48. 48.Zeroual, I., Goldhahn, D., Eckart, T., and Lakhouaja, A. (2019). OSIAN: Open source international Arabic news corpus - preparation and integration into the CLARIN-infrastructure. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pages 175–182, Florence, Italy, August. Association for Computational Linguistics.
  49. 49.Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S. (2015). Aligning books and movies: Towards story-like visual explanations by watching movies and reading books.

Citation

MLA
Antoun, W., et al. “AraBERT: Transformer-based Model for Arabic Language Understanding”. arXiv, 2020, http://arxiv.org/abs/2003.00104v4.
APA
Antoun, W., Baly, F., & Hajj, H. (2020). AraBERT: Transformer-based Model for Arabic Language Understanding. arXiv. http://arxiv.org/abs/2003.00104v4
Chicago
Antoun, W., F. Baly, and H. Hajj. 2020. “AraBERT: Transformer-based Model for Arabic Language Understanding”. arXiv. http://arxiv.org/abs/2003.00104v4.
Harvard
Antoun, W., Baly, F. and Hajj, H. (2020) “AraBERT: Transformer-based Model for Arabic Language Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2003.00104v4.
Vancouver
1. Antoun W, Baly F, Hajj H (2020) AraBERT: Transformer-based Model for Arabic Language Understanding. arXiv

BibTeX

@article{antoun2020arabert,
  title = {AraBERT: Transformer-based Model for Arabic Language Understanding},
  author = {Antoun, Wissam and Baly, Fady and Hajj, Hazem},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2003.00104v4},
  eprint = {2003.00104}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF