AceGPT, Localizing Large Language Models in Arabic
Huang HuangFei YuJianqing ZhuXuening SunHao ChengDingjie SongZhihong ChenMosen AlharthiBang AnJuncai He
Presents AceGPT, an open-source Arabic large language model that addresses cultural alignment through localized pre-training, native instruction tuning, and culturally tuned reinforcement learning to achieve state-of-the-art performance across Arabic benchmarks.
Mainstream artificial intelligence models frequently struggle with cultural sensitivity and regional alignment when deployed outside Western contexts. In Arabic-speaking communities, existing open-source and proprietary language models often reflect English-centric cultural biases, primarily because their training data heavily relies on direct translations of Western datasets. This reliance undermines the models' ability to reflect local values, literary traditions, and regional perspectives, creating substantial compliance, social, and operational risks for practical adoption in the Arab world.
The article demonstrates an end-to-end localization framework designed to develop AceGPT, an open-source Arabic-centric large language model family available in 7-billion and 13-billion parameter configurations. The work evaluates whether systematically adapting pre-training, fine-tuning, and reward alignment processes can effectively bridge cultural gaps while preserving competitive core language understanding and reasoning performance.
The authors implemented a three-stage training pipeline. First, base models were further pre-trained on targeted corpora comprising 10 to 30 billion tokens across Arabic and English to reinforce linguistic fundamentals. Second, the supervised fine-tuning stage replaced translated Western queries with localized real-world Arabic questions, pairing them with native GPT-4 responses generated directly within an Arabic cultural framing. Third, the authors applied reinforcement learning from artificial intelligence feedback, training a culturally attuned reward model using paired preference judgments to align conversational outputs directly with Arab cultural norms, laws, and customs. Performance was comprehensively assessed across multiple benchmarks, including traditional natural language understanding tasks, instruction-following tests, knowledge examinations, and a newly developed 8,000-question Arabic Cultural and Value Alignment benchmark.
The key findings show significant improvements across cultural and functional capabilities. AceGPT models established state-of-the-art results among open-source Arabic models on the instruction-following benchmarks, Arabic Vicuna-80 and Arabic AlpacaEval, achieving roughly 30% to 33% higher performance relative to leading open baselines. On cultural and value alignment evaluations, AceGPT-13B achieved a 75.02% F1 score, outperforming open alternatives and trailing commercial proprietary systems like GPT-3.5 Turbo by only 0.55%. In blind evaluations by native Arabic speakers, AceGPT-13B consistently surpassed competing open models and attained a 69.7% win-or-tie rate against GPT-3.5 Turbo on Arabic AlpacaEval. Furthermore, ablation analyses confirmed that applying reinforcement learning with localized feedback drove substantial localization improvements, boosting cultural alignment scores by up to 27% on the 7-billion model while improving general response quality.
These results demonstrate that direct translation is insufficient for building regional artificial intelligence solutions, whereas end-to-end cultural adaptation provides a cost-effective pathway to close the performance gap with proprietary models. Aligning language models directly with native sociocultural norms mitigates reputational and compliance risks in enterprise deployments while improving user trust and interaction quality. However, incorporating uncurated Western dialogue data during fine-tuning was shown to degrade cultural alignment, highlighting the need for strict data selection governance.
Organizations planning regional language model deployments should adopt end-to-end localized fine-tuning and native preference alignment rather than relying on translated instruction sets. Before enterprise deployment, teams must conduct additional testing on safety, reasoning, and misinformation, as these dimensions were outside the scope of the study. Further research and engineering should focus on expanding the model's vocabulary encoding efficiency, scaling Arabic training tokens, and enriching cultural evaluation datasets to ensure robust real-world reliability.
- Paper: Llama 2: Open Foundation and Fine-Tuned Chat Models, Hugo Touvron et al. (2023). AceGPT adapts the Llama 2 foundation model, so this paper clarifies the base model and alignment methods it builds upon.
- Paper: LLaMA: Open and Efficient Foundation Language Models, Hugo Touvron et al. (2023). Reading about LLaMA first traces the open foundation-model lineage that leads to the Llama 2 base AceGPT localizes.
No sufficiently relevant recommendations were found.
