AceGPT, Localizing Large Language Models in Arabic

Huang HuangFei YuJianqing ZhuXuening SunHao ChengDingjie SongZhihong ChenMosen AlharthiBang AnJuncai He

article2024NAACL137 citations

Presents AceGPT, an open-source Arabic large language model that addresses cultural alignment through localized pre-training, native instruction tuning, and culturally tuned reinforcement learning to achieve state-of-the-art performance across Arabic benchmarks.

Listen

Mainstream artificial intelligence models frequently struggle with cultural sensitivity and regional alignment when deployed outside Western contexts. In Arabic-speaking communities, existing open-source and proprietary language models often reflect English-centric cultural biases, primarily because their training data heavily relies on direct translations of Western datasets. This reliance undermines the models' ability to reflect local values, literary traditions, and regional perspectives, creating substantial compliance, social, and operational risks for practical adoption in the Arab world.

The article demonstrates an end-to-end localization framework designed to develop AceGPT, an open-source Arabic-centric large language model family available in 7-billion and 13-billion parameter configurations. The work evaluates whether systematically adapting pre-training, fine-tuning, and reward alignment processes can effectively bridge cultural gaps while preserving competitive core language understanding and reasoning performance.

The authors implemented a three-stage training pipeline. First, base models were further pre-trained on targeted corpora comprising 10 to 30 billion tokens across Arabic and English to reinforce linguistic fundamentals. Second, the supervised fine-tuning stage replaced translated Western queries with localized real-world Arabic questions, pairing them with native GPT-4 responses generated directly within an Arabic cultural framing. Third, the authors applied reinforcement learning from artificial intelligence feedback, training a culturally attuned reward model using paired preference judgments to align conversational outputs directly with Arab cultural norms, laws, and customs. Performance was comprehensively assessed across multiple benchmarks, including traditional natural language understanding tasks, instruction-following tests, knowledge examinations, and a newly developed 8,000-question Arabic Cultural and Value Alignment benchmark.

The key findings show significant improvements across cultural and functional capabilities. AceGPT models established state-of-the-art results among open-source Arabic models on the instruction-following benchmarks, Arabic Vicuna-80 and Arabic AlpacaEval, achieving roughly 30% to 33% higher performance relative to leading open baselines. On cultural and value alignment evaluations, AceGPT-13B achieved a 75.02% F1 score, outperforming open alternatives and trailing commercial proprietary systems like GPT-3.5 Turbo by only 0.55%. In blind evaluations by native Arabic speakers, AceGPT-13B consistently surpassed competing open models and attained a 69.7% win-or-tie rate against GPT-3.5 Turbo on Arabic AlpacaEval. Furthermore, ablation analyses confirmed that applying reinforcement learning with localized feedback drove substantial localization improvements, boosting cultural alignment scores by up to 27% on the 7-billion model while improving general response quality.

These results demonstrate that direct translation is insufficient for building regional artificial intelligence solutions, whereas end-to-end cultural adaptation provides a cost-effective pathway to close the performance gap with proprietary models. Aligning language models directly with native sociocultural norms mitigates reputational and compliance risks in enterprise deployments while improving user trust and interaction quality. However, incorporating uncurated Western dialogue data during fine-tuning was shown to degrade cultural alignment, highlighting the need for strict data selection governance.

Organizations planning regional language model deployments should adopt end-to-end localized fine-tuning and native preference alignment rather than relying on translated instruction sets. Before enterprise deployment, teams must conduct additional testing on safety, reasoning, and misinformation, as these dimensions were outside the scope of the study. Further research and engineering should focus on expanding the model's vocabulary encoding efficiency, scaling Arabic training tokens, and enriching cultural evaluation datasets to ensure robust real-world reliability.

No sufficiently relevant recommendations were found.

Cover for AceGPT, Localizing Large Language Models in Arabic

Abstract

This paper is devoted to the development of a localized Large Language Model (LLM) specifically for Arabic, a language imbued with unique cultural characteristics inadequately addressed by current mainstream models. Significant concerns emerge when addressing cultural sensitivity and local values. To address this, the paper proposes a comprehensive solution that includes further pre-training with Arabic texts, Supervised Fine-Tuning (SFT) utilizing native Arabic instructions, and GPT-4 responses in Arabic, alongside Reinforcement Learning with AI Feedback (RLAIF) employing a reward model attuned to local culture and values. The goal is to cultivate culturally cognizant and value-aligned Arabic LLMs capable of accommodating the diverse, application-specific needs of Arabic-speaking communities. Comprehensive evaluations reveal that the resulting model, dubbed ‘AceGPT’, sets the state-of-the-art standard for open Arabic LLMs across various benchmarks. Codes, data, and models are in https://github.com/FreedomIntelligence/AceGPT.

Table of Contents

  • 1 Introduction
  • 2 Recipe of AceGPT
  • 2.1 Motivation: The Localization Issue
  • 2.2 Methodology of AceGPT
  • 2.2.1 Localized Pre-Training
  • 2.2.2 Localized Supervised Fine-Tuning
  • 2.2.3 Reinforcement Learning from AI Feedback
  • 3 Localization Evaluation
  • 3.1 Evaluation Protocol
  • 3.2 Experiment Results
  • 4 Overall Evaluation
  • 4.1 Evaluation Protocol
  • 4.2 Experiment Results
  • 5 Experimental Analysis
  • 5.1 On Pre-Training
  • 5.2 On Supervised Fine-Tuning
  • 5.3 On RLAIF
  • 5.3.1 Reward Model
  • 5.3.2 Ablation
  • 6 Conclusion
  • Limitation
  • Ethical Statement
  • Acknowledgement
  • References
  • A Related Work
  • B Localization Issues
  • B.1 Sample Questions for Localization
  • B.2 Case Study
  • C Construction of ACVA
  • D Preference Data for RLAIF
  • E Implementation of Training
  • E.1 Pre-Training
  • E.2 Supervised Fine-Tuning
  • E.3 Reward Model Training
  • E.4 PPO
  • F Implementation of Evaluation
  • F.1 Baselines and Benchmarks
  • F.2 Evaluation on Instruction Following
  • F.3 Evaluation on Knowledge
  • F.4 Evaluation on ACVA
  • G More Experiments of AceGPT Evaluation
  • G.1 Supplementary Experimental Results
  • G.2 Evaluation on Arabic NLU Tasks
  • H Relationship between Arabic Culture and Arabic Language

Knowls

  1. Knowl 1 — AceGPT’s three-stage Arabic localization pipeline

    model/method

    AceGPT adapts LLaMA 2 for Arabic through three stages: continued pre-training on Arabic-rich text to improve Arabic language and cultural coverage; supervised fine-tuning on Arabic instructions and responses to support instruction following; and reinforcement learning from AI feedback (RLAIF) using preferences intended to reflect Arabic cultural and value norms. The continued-pre-training models are named AceGPT-base, while the conversational models produced through supervised fine-tuning and RLAIF are named AceGPT-chat. The paper develops 7B- and 13B-parameter versions.

  2. Knowl 2 — Arabic-rich continued pre-training of LLaMA 2

    model/method

    AceGPT-base is initialized from LLaMA 2 and further pre-trained on mixed Arabic and English text, with more tokens allocated to Arabic. The Arabic corpus combines ArabicText-2022 with material refined from Arabic Wikipedia, CC100, and OSCAR3; English text is sampled from SlimPajama to help retain English knowledge. The 7B model is trained on 30B tokens—19.2B Arabic and 10.8B English—while the 13B model is trained on 10B tokens—6B Arabic and 4B English. The authors retain the LLaMA 2 vocabulary, which they report contains all 53 Arabic letters, rather than expanding it to reduce training costs. Training uses a 2048-token context, AdamW, cosine learning-rate scheduling, a learning rate of 10−410^{-4}, gradient accumulation of 128, total batch size 3072, and a warm-up comprising 5% of training.

  3. Knowl 3 — Arabic-native supervised instruction fine-tuning

    model/method

    AceGPT-chat is supervised-fine-tuned on Arabic instructions and responses designed to reflect real-world Arabic-language use. The central localized source is Arabic Quora questions, with GPT-4-generated Arabic answers. The training mixture also includes Alpaca, Evol-Instruct, and Code-Alpaca: their English instructions are translated into Arabic and their answers are regenerated by GPT-4, rather than translated from existing answers. Original ShareGPT conversations are retained because regenerating their responses would change the dialogues. The authors report a 629,293-example mixture trained for one epoch; native Arabic Quora and Alpaca-Arabic data are included three times, while ShareGPT and Alpaca-Chinese are included once to limit the non-Arabic proportion. The 7B and 13B models are fine-tuned on eight A100 80GB GPUs, using maximum learning rates of 5×10−55\times10^{-5} and 1×10−51\times10^{-5}, respectively, with AdamW, cosine scheduling, and a 0.03 warm-up rate.

  4. Knowl 4 — Localized preference modeling and RLAIF

    model/method

    AceGPT’s RLAIF stage uses a reward model trained on Arabic-localized preferences and then optimizes the chat model with PPO. For Arabic preference collection, the authors reuse 40K Quora instructions, sample paired responses from the supervised-fine-tuned 7B model, and ask GPT-4 to choose the better response. Because GPT-4 showed position bias, they reverse the response order and retain preferences consistent across both runs, yielding 12K Arabic preference pairs. A study of 800 examples found a correlation of 0.84 between GPT-4 and human preference judgments. The reward-model training also includes 12K examples sampled from public English preference datasets for generalization. The 7B reward model is initialized from Ziya and scores a chosen response ycy_c above a rejected response yry_r for input xx using the loss

    L(θ)=−E(x,yc,yr)∼D[log⁡σ(rθ(x,yc)−rθ(x,yr))],\mathcal{L}(\theta)=-\mathbb{E}_{(x,y_c,y_r)\sim D}\left[\log\sigma\left(r_\theta(x,y_c)-r_\theta(x,y_r)\right)\right],

    where DD is the preference-pair dataset, rθr_\theta is the scalar reward function with parameters θ\theta, and σ\sigma is the logistic sigmoid. PPO uses 30K additional Quora questions, distinct from those used for preference collection. Its clipped policy objective is Et[min⁡(ρtAt,clip⁡(ρt,1−ϵ,1+ϵ)At)]\mathbb{E}_t[\min(\rho_t A_t,\operatorname{clip}(\rho_t,1-\epsilon,1+\epsilon)A_t)], where ρt\rho_t is the current-to-old policy probability ratio for the next token, AtA_t is its advantage, and ϵ\epsilon is the clipping parameter. The reported PPO setup uses a KL penalty of 0.01, policy clipping threshold 0.2, value-loss clipping threshold 0.3, reward clipping to [−5,5][-5,5], and generalized-advantage-estimation parameters γ=1\gamma=1 and λ=0.95\lambda=0.95. Both 7B and 13B policies use the same 7B reward model, trained on preferences from the 7B policy.

  5. Knowl 5 — Arabic Cultural and Value Alignment benchmark

    experimental setup

    The Arabic Cultural and Value Alignment (ACVA) benchmark tests cultural and value alignment with Arabic yes/no questions. The authors gathered more than 50 topics spanning countries, civilization relations, science and humanity, manners, and religion, then used GPT-3.5 Turbo to generate over 8,000 Arabic questions with balanced true and false labels. Arabic speakers reviewed questions and labels for a sample covering 50% of the topics; this review produced a 2,486-example Clean set, alongside the larger All set. The authors evaluate fine-tuned chat models zero-shot using F1. They report a Pearson correlation of 0.9825 between accuracy on the All and Clean sets.

  6. Knowl 6 — Localization evaluations show stronger Arabic-cultural alignment

    empirical result

    On zero-shot ACVA, AceGPT-7B-chat scores 69.53% F1 on the All set and 70.03% on the Clean set; AceGPT-13B-chat scores 75.02% and 74.62%, respectively. Jais-13B-chat scores 61.44% and 66.83%, while GPT-3.5 Turbo scores 75.57% and 79.03%. Thus both AceGPT sizes outperform the compared open-source chat models, and AceGPT-13B-chat is close to GPT-3.5 Turbo on the All set. A separate audit of responses to 20 Arabic questions counted Arabic entities: for people, Arabic names accounted for 3/25 (12.00%) in Jais-13B, 12/45 (26.67%) in GPT-3.5 Turbo, 22/56 (39.29%) in GPT-4, and 31/62 (50.00%) in AceGPT; for locations, the corresponding proportions were 3/16 (18.75%), 13/48 (27.08%), 16/74 (21.62%), and 11/38 (28.95%). The entity audit and ACVA scores provide distinct evidence about localization: the audit measures Arabic-rooted entities in generated answers, while ACVA measures agreement with labeled cultural and value statements.

  7. Knowl 7 — Instruction-following performance against open Arabic chat models

    empirical result

    On Arabic Vicuna-80 and Arabic AlpacaEval, GPT-4 scored model responses against GPT-3.5 Turbo, whose performance ratio is 100%. AceGPT-7B-chat reaches 94.82% ± 0.2 on Vicuna-80 and 93.81% ± 0.1 on AlpacaEval; AceGPT-13B-chat reaches 100.88% ± 0.4 and 97.95% ± 0.1. In human evaluations by three native Arabic speakers, AceGPT-7B-chat is preferred or tied over Jais-13B-chat in 89.2% of Vicuna-80 comparisons and 89.5% of AlpacaEval comparisons; for AceGPT-13B-chat, those figures are 89.6% and 92.2%. Against GPT-3.5 Turbo, AceGPT-7B-chat wins or ties in 60.4% of Vicuna-80 and 66.7% of AlpacaEval comparisons, while AceGPT-13B-chat wins or ties in 73.4% and 69.7%. The human results therefore show an advantage over Jais but do not establish that AceGPT consistently outperforms GPT-3.5 Turbo.

  8. Knowl 8 — Knowledge and Arabic NLU results

    empirical result

    In few-shot Arabic MMLU, AceGPT-13B-base achieves 40.45% average accuracy, above the compared open-source models, including Jais-30B-base at 36.27%; GPT-3.5 Turbo scores 49.07%. AceGPT-7B-base reaches 32.14% average. On EXAMs, AceGPT-13B-base scores 36.63%, below Jais-30B-base at 39.91% and GPT-3.5 Turbo at 45.63%. On the nine-task ALUE Arabic language-understanding benchmark, AceGPT-13B-base obtains a reported average score of 72.8, ranking behind AraMUS at 74.0. These results indicate strong open-model performance in Arabic MMLU and ALUE, but not a lead on every knowledge evaluation.

  9. Knowl 9 — Ablations attribute localization gains to pre-training, Quora data, and RLAIF

    empirical result

    Several ablations test the contributions of the localization pipeline. In few-shot ACVA evaluation, continued pre-training raises the 7B model’s F1 from 51.44% for LLaMA 2 to 68.28% for AceGPT-base, and raises the 13B model’s F1 from 65.67% to 76.23%. In the supervised-data comparison, adding Quora after the reported Alpaca-Arabic, ShareGPT, and Evol-Instruct configurations raises ACVA F1 from 61.72% to 65.53%; Evol-Instruct produces the highest instruction-following ratios among those configurations, while ShareGPT is associated with lower ACVA F1. Adding RLAIF raises 7B ACVA F1 from 42.48% to 69.53%, Arabic Vicuna-80 from 92.01% to 94.82%, and Arabic AlpacaEval from 91.35% to 93.81%. For 13B, the corresponding changes are 74.18% to 75.02%, 95.14% to 100.88%, and 93.05% to 97.95%. The reported RLAIF gains are especially large for 7B on ACVA; the authors suggest that the 7B preference data and reward model may fit that policy better, and that 13B may have less room for improvement.

  10. Knowl 10 — Stated limitations of AceGPT

    limitation

    The authors identify several limitations. AceGPT retains the LLaMA 2 vocabulary rather than expanding it, which they say affects Arabic text-encoding efficiency; limited compute also restricted the number of Arabic training tokens. The evaluation omits reasoning, misinformation, and bias testing, leaving safety alignment insufficiently assessed and leading the authors to limit the model’s intended use to academic research rather than online deployment. They also state that the cultural dataset needs improvements in both quality and quantity, despite manual checks.

Coverage note — Detailed ALUE per-task scores, evaluation prompt templates, and illustrative response case studies are omitted as supplementary protocol or example material; the main benchmark outcome and the paper’s central localization evidence are retained.

References

  1. 1.Asaad Alghamdi, Xinyu Duan, Wei Jiang, Zhenhai Wang, Yimeng Wu, Qingrong Xia, Zhefeng Wang, Yi Zheng, Mehdi Rezagholizadeh, Baoxing Huai, et al. 2023. Aramus: Pushing the limits of data and model scale for arabic natural language processing. arXiv preprint arXiv:2306.06800.
  2. 2.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. CoRR, abs/2204.05862.
  3. 3.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712.
  4. 4.Ewa Callahan and Susan C. Herring. 2011. Cultural bias in wikipedia content on famous persons. J. Assoc. Inf. Sci. Technol., 62(10):1899–1915.
  5. 5.Sky CH-Wang, Arkadiy Saakyan, Oliver Li, Zhou Yu, and Smaranda Muresan. 2023. Sociocultural norm similarities and differences via situational alignment and explainable textual entailment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 3548–3564. Association for Computational Linguistics.
  6. 6.Sahil Chaudhary. 2023. Code alpaca: An instruction-following llama model for code generation. https://github.com/sahil280114/codealpaca.
  7. 7.Zhihong Chen, Feng Jiang, Junying Chen, Tiannan Wang, Fei Yu, Guiming Chen, Hongbo Zhang, Juhao Liang, Chen Zhang, Zhiyi Zhang, et al. 2023a. Phoenix: Democratizing chatgpt across languages. arXiv preprint arXiv:2304.10453.
  8. 8.Zhihong Chen, Shuo Yan, Juhao Liang, Feng Jiang, Xiangbo Wu, Fei Yu, Guiming Hardy Chen, Junying Chen, Hongbo Zhang, Li Jianquan, Wan Xiang, and Benyou Wang. 2023b. MultilingualSIFT: Multilingual Supervised Instruction Fine-tuning.
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  10. 10.Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. Xnli: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2475–2485.
  11. 11.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  12. 12.Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 1286–1305. Association for Computational Linguistics.
  13. 13.Yann Dubois, Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpaca-farm: A simulation framework for methods that learn from human feedback.
  14. 14.Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. 2020. EXAMS: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 5427–5444. Association for Computational Linguistics.
  15. 15.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  16. 16.Daniel Hershcovich, Stella Frank, Heather C. Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and Anders Søgaard. 2022. Challenges and strategies in cross-cultural NLP. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 6997–7013. Association for Computational Linguistics.
  17. 17.Jing Huang and Diyi Yang. 2023. Culturally aware natural language inference. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 7591–7609. Association for Computational Linguistics.
  18. 18.Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.
  19. 19.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. 2023. Openassistant conversations - democratizing large language model alignment. CoRR, abs/2304.07327.
  20. 20.Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi. 2023. RLAIF: scaling reinforcement learning from human feedback with AI feedback. CoRR, abs/2309.00267.
  21. 21.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. 2022. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786.
  22. 22.Tarek Naous, Michael J. Ryan, and Wei Xu. 2023. Having beer after prayer? measuring cultural bias in large language models. CoRR, abs/2305.14456.
  23. 23.Shramay Palta and Rachel Rudinger. 2023. FORK: A bite-sized test set for probing culinary cultural biases in commonsense reasoning models. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, pages 9952–9962. Association for Computational Linguistics.
  24. 24.Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao. 2023. Instruction tuning with gpt-4. arXiv preprint arXiv:2304.03277.
  25. 25.Aida Ramezani and Yang Xu. 2023. Knowledge of cultural moral norms in large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 428–446. Association for Computational Linguistics.
  26. 26.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. CoRR, abs/1707.06347.
  27. 27.Haitham Seelawi, Ibraheem Tuffaha, Mahmoud Gzawi, Wael Farhan, Bashar Talafha, Riham Badawi, Zyad Sober, Oday Al-Dweik, Abed Alhakim Freihat, and Hussein Al-Natsheh. 2021. ALUE: Arabic language understanding evaluation. In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 173–184, Kyiv, Ukraine (Virtual). Association for Computational Linguistics.
  28. 28.Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, Osama Mohammed Afzal, Samta Kamboj, Onkar Pandit, Rahul Pal, et al. 2023. Jais and jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models. arXiv preprint arXiv:2308.16149.
  29. 29.Daria Soboleva, Al-Khateeb Faisal, Myers Robert Steeves Jacob R, Hestness Joel, and Dey Nolan. 2023. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. www.cerebras.net/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama.
  30. 30.Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. 2020. Learning to summarize from human feedback. CoRR, abs/2009.01325.
  31. 31.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  32. 32.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  33. 33.Peiyi Wang, Lei Li, Liang Chen, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023a. Large language models are not fair evaluators. CoRR, abs/2305.17926.
  34. 34.Wenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai, Jen-tse Huang, Zhaopeng Tu, and Michael R. Lyu. 2023b. Not all countries celebrate thanksgiving: On the cultural dominance in large language models. CoRR, abs/2310.12481.
  35. 35.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen. 2022. Super-naturalinstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 5085–5109. Association for Computational Linguistics.
  36. 36.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244.
  37. 37.Hongbo Zhang, Junying Chen, Feng Jiang, Fei Yu, Zhihong Chen, Jianquan Li, Guiming Chen, Xiangbo Wu, Zhiyi Zhang, Qingying Xiao, et al. 2023. Huatuogpt, towards taming language model to be a doctor. arXiv preprint arXiv:2305.15075.

Citation

MLA
Huang, H., et al. “AceGPT, Localizing Large Language Models in Arabic”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 8139–63, https://doi.org/10.18653/v1/2024.naacl-long.450.
APA
Huang, H., Yu, F., Zhu, J., Sun, X., Cheng, H., Dingjie, S., Chen, Z., Alharthi, M., An, B., He, J., Liu, Z., Chen, J., Li, J., Wang, B., Zhang, L., Sun, R., Wan, X., Li, H., & Xu, J. (2024). AceGPT, Localizing Large Language Models in Arabic. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 8139–8163. https://doi.org/10.18653/v1/2024.naacl-long.450
Chicago
Huang, H., F. Yu, J. Zhu, et al. 2024. “AceGPT, Localizing Large Language Models in Arabic”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 8139–63. https://doi.org/10.18653/v1/2024.naacl-long.450.
Harvard
Huang, H. et al. (2024) “AceGPT, Localizing Large Language Models in Arabic”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8139–8163. Available at: https://doi.org/10.18653/v1/2024.naacl-long.450.
Vancouver
1. Huang H, Yu F, Zhu J, et al (2024) AceGPT, Localizing Large Language Models in Arabic. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 8139–8163

BibTeX

@inproceedings{huang-etal-2024-acegpt,
    title = "{A}ce{GPT}, Localizing Large Language Models in {A}rabic",
    author = "Huang, Huang  and
      Yu, Fei  and
      Zhu, Jianqing  and
      Sun, Xuening  and
      Cheng, Hao  and
      Dingjie, Song  and
      Chen, Zhihong  and
      Alharthi, Mosen  and
      An, Bang  and
      He, Juncai  and
      Liu, Ziche  and
      Chen, Junying  and
      Li, Jianquan  and
      Wang, Benyou  and
      Zhang, Lian  and
      Sun, Ruoyu  and
      Wan, Xiang  and
      Li, Haizhou  and
      Xu, Jinchao",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.450/",
    doi = "10.18653/v1/2024.naacl-long.450",
    pages = "8139--8163"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/