XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
Dasol ChoiEugenia KimJae-won NohSanghyun SeoEunmi KimYunjin ParkBrigitta Jesica KartonoJosef PichlmeierHelena BerndtSai Krishna Mendu
Presents XL-SafetyBench, a cross-cultural benchmark spanning 10 country-language pairs that separates adversarial attack defense from cultural sensitivity, revealing that frontier models decouple safety from local awareness while local models frequently mistake generation failure for alignment.
As artificial intelligence systems are deployed globally, existing safety evaluations remain overwhelmingly English-centric and reliant on translated prompts. This standard approach fails to capture how safety risks and social harms natively manifest across distinct legal, economic, and cultural environments. Furthermore, conventional benchmarks typically measure safety as a single, uniform dimension, overlooking whether models can recognize culturally embedded taboos in everyday contexts.
The article introduces and evaluates XL-SafetyBench, a benchmark designed to assess large language models across two distinct safety dimensions: adversarial robustness against country-specific harms and cultural sensitivity awareness within natural tasks. The evaluation covers 10 country-language pairs spanning North America, Europe, Asia, and the Middle East: the United States, France, Germany, Spain, South Korea, Japan, India, Indonesia, Türkiye, and the United Arab Emirates.
The benchmark comprises 5,500 expert-validated test cases generated through automated discovery, red-teaming pipelines, and multi-stage human validation by 20 native-speaker experts. It evaluates 10 global frontier models and 27 country-specific local models across two core tracks. The Jailbreak Benchmark tests model resistance against localized adversarial prompts (measuring the Attack Success Rate and Neutral-Safe Rate), while the Cultural Benchmark tests whether models detect subtle cultural taboos embedded inside innocuous, everyday requests (measuring the Cultural Sensitivity Rate).
The evaluation reveals several critical findings regarding current model capabilities. First, adversarial safety and cultural sensitivity are largely uncoupled among leading global models; high resistance to adversarial attacks does not reliably predict an awareness of local cultural norms. Second, frontier models exhibit strong geographic disparities, performing best on United States prompts (34.5% attack success rate, 69.5% cultural sensitivity rate) while showing substantially higher vulnerabilities in markets such as the UAE and South Korea, alongside steep drops in cultural awareness in India and Türkiye. Third, open-weight global models perform poorly, with jailbreak success rates exceeding 90% and cultural sensitivity rates below 15%. Finally, many country-specific local models display an illusion of safety driven by an inverse trade-off between attack success and neutral-safe rates (r = -0.81); their low attack rates frequently stem from general comprehension failures and incoherent outputs rather than principled safety alignment.
These findings indicate that deploying AI models globally based solely on aggregate or English-centric safety scores creates substantial operational, compliance, and reputational risks. The inability of models to navigate localized legal landmines or cultural norms can lead to severe real-world fallout. Furthermore, local language pre-training alone does not guarantee cultural competence, as even the largest local models struggle to match the nuanced reasoning of leading systems.
Organizations deploying AI internationally should avoid using single composite safety scores and instead evaluate adversarial robustness and cultural competence as separate metrics. Developers must incorporate culturally grounded alignment workflows rather than relying on translation-based safety data or simple local-language pre-training. Further research is necessary to expand benchmark coverage to multilingual nations, capture regional variations across culturally distinct countries sharing a language (such as Spain and Latin American nations), and address hardware-level inference limits observed in smaller models.
The benchmark results provide a high-confidence assessment of country-level safety performance, supported by substantial agreement between native-speaker human annotators and automated judges. However, stakeholders should interpret granular category-level metrics with caution, as the sample size of 100 cultural scenarios per country is optimized for country-level comparisons rather than sub-category statistical power.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). HarmBench provides the foundational standardized evaluation framework for automated red teaming and attack success rate measurement that XL-SafetyBench adapts and extends into multilingual, culturally grounded settings.
- Paper: Unintended Impacts of LLM Alignment on Global Representation, Michael J. Ryan et al. (2024). This study analyzes how standard alignment pipelines create geographic and cultural disparities across languages and national opinions, motivating XL-SafetyBench's explicit decoupling of cultural sensitivity from universal harm alignment.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This work establishes the concepts of competing objectives and mismatched generalization in safety training, explaining why cross-cultural and non-English adversarial prompts reliably bypass standard guardrails.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). TrustLLM defines multidimensional trustworthiness taxonomies and documents the exaggerated safety and refusal phenomena that XL-SafetyBench addresses via its Neutral-Safe Rate.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). MEGAVERSE establishes broad cross-lingual benchmarking methodologies and demonstrates the severe performance drops of LLMs on non-English inputs that inform multilingual safety evaluations.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). Llama Guard introduces the standard input-output moderation taxonomy and classifier architectures against which modern adversarial and cultural safety benchmarks measure model responses.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). This foundational paper introduces automated language model red teaming, establishing the LLM-assisted prompt generation pipelines built upon by XL-SafetyBench.
- Paper: AlignBench: Benchmarking Chinese Alignment of Large Language Models, Xiao Liu et al. (2024). AlignBench demonstrates the necessity of moving beyond translated English benchmarks by creating native, culturally specific alignment evaluations in Chinese.
No sufficiently relevant recommendations were found.
