XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity

Dasol ChoiEugenia KimJae-won NohSanghyun SeoEunmi KimYunjin ParkBrigitta Jesica KartonoJosef PichlmeierHelena BerndtSai Krishna Mendu

article2026arXiv3 citations

Presents XL-SafetyBench, a cross-cultural benchmark spanning 10 country-language pairs that separates adversarial attack defense from cultural sensitivity, revealing that frontier models decouple safety from local awareness while local models frequently mistake generation failure for alignment.

Listen

As artificial intelligence systems are deployed globally, existing safety evaluations remain overwhelmingly English-centric and reliant on translated prompts. This standard approach fails to capture how safety risks and social harms natively manifest across distinct legal, economic, and cultural environments. Furthermore, conventional benchmarks typically measure safety as a single, uniform dimension, overlooking whether models can recognize culturally embedded taboos in everyday contexts.

The article introduces and evaluates XL-SafetyBench, a benchmark designed to assess large language models across two distinct safety dimensions: adversarial robustness against country-specific harms and cultural sensitivity awareness within natural tasks. The evaluation covers 10 country-language pairs spanning North America, Europe, Asia, and the Middle East: the United States, France, Germany, Spain, South Korea, Japan, India, Indonesia, Türkiye, and the United Arab Emirates.

The benchmark comprises 5,500 expert-validated test cases generated through automated discovery, red-teaming pipelines, and multi-stage human validation by 20 native-speaker experts. It evaluates 10 global frontier models and 27 country-specific local models across two core tracks. The Jailbreak Benchmark tests model resistance against localized adversarial prompts (measuring the Attack Success Rate and Neutral-Safe Rate), while the Cultural Benchmark tests whether models detect subtle cultural taboos embedded inside innocuous, everyday requests (measuring the Cultural Sensitivity Rate).

The evaluation reveals several critical findings regarding current model capabilities. First, adversarial safety and cultural sensitivity are largely uncoupled among leading global models; high resistance to adversarial attacks does not reliably predict an awareness of local cultural norms. Second, frontier models exhibit strong geographic disparities, performing best on United States prompts (34.5% attack success rate, 69.5% cultural sensitivity rate) while showing substantially higher vulnerabilities in markets such as the UAE and South Korea, alongside steep drops in cultural awareness in India and Türkiye. Third, open-weight global models perform poorly, with jailbreak success rates exceeding 90% and cultural sensitivity rates below 15%. Finally, many country-specific local models display an illusion of safety driven by an inverse trade-off between attack success and neutral-safe rates (r = -0.81); their low attack rates frequently stem from general comprehension failures and incoherent outputs rather than principled safety alignment.

These findings indicate that deploying AI models globally based solely on aggregate or English-centric safety scores creates substantial operational, compliance, and reputational risks. The inability of models to navigate localized legal landmines or cultural norms can lead to severe real-world fallout. Furthermore, local language pre-training alone does not guarantee cultural competence, as even the largest local models struggle to match the nuanced reasoning of leading systems.

Organizations deploying AI internationally should avoid using single composite safety scores and instead evaluate adversarial robustness and cultural competence as separate metrics. Developers must incorporate culturally grounded alignment workflows rather than relying on translation-based safety data or simple local-language pre-training. Further research is necessary to expand benchmark coverage to multilingual nations, capture regional variations across culturally distinct countries sharing a language (such as Spain and Latin American nations), and address hardware-level inference limits observed in smaller models.

The benchmark results provide a high-confidence assessment of country-level safety performance, supported by substantial agreement between native-speaker human annotators and automated judges. However, stakeholders should interpret granular category-level metrics with caution, as the sample size of 100 cultural scenarios per country is optimized for country-level comparisons rather than sub-category statistical power.

No sufficiently relevant recommendations were found.

Cover for XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity

Abstract

Current LLM safety benchmarks are predominantly English-centric and often rely on translation, failing to capture country-specific harms. Moreover, they rarely evaluate a model's ability to detect culturally embedded sensitivities as distinct from universal harms. We introduce XL-SafetyBench. a suite of 5,500 test cases across 10 country-language pairs, comprising a Jailbreak Benchmark of country-grounded adversarial prompts and a Cultural Benchmark where local sensitivities are embedded within innocuous requests. Each item is constructed via a multi-stage pipeline that combines LLM-assisted discovery, automated validation gates, and dual independent native-speaker annotators per country. To distinguish principled refusal from comprehension failure, we evaluate Attack Success Rate (ASR) alongside two complementary metrics we introduce: Neutral-Safe Rate (NSR) and Cultural Sensitivity Rate (CSR). Evaluating 10 frontier and 27 local LLMs reveals two key findings. First, jailbreak robustness and cultural awareness do not show a coupled relationship among frontier models, so a composite safety score obscures per-axis variation. Second, local models exhibit a near-linear ASR-NSR trade-off (r = -0.81), indicating that their apparent safety reflects generation failure rather than genuine alignment. XL-SafetyBench enables more nuanced, cross-cultural safety evaluation in the multilingual era.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Multilingual safety benchmarks
  • 2.2 Cultural knowledge evaluation in LLMs
  • 3 The XL-SafetyBench Framework
  • 3.1 Jailbreak Benchmark: Country-Specific Adversarial Robustness
  • 3.2 Cultural Benchmark: Culturally Embedded Sensitivities
  • 3.3 Evaluation Metric
  • 4 Experimental Setup
  • 5 Results and Analysis
  • 5.1 Global Model Performance
  • 5.2 Country-Specific Local Models
  • 6 Limitations, Future Work, and Broader Impacts
  • 7 Conclusion
  • References
  • Appendix Contents
  • A Implementation Details
  • A.1 LLMs Used in Dataset Construction
  • A.2 Inference Settings
  • B Dataset Generation Prompts
  • B.1 Jailbreak Benchmark: Subcategory Generation
  • B.2 Jailbreak Benchmark: Base Query Generation
  • B.3 Jailbreak Benchmark: Attack Prompt Generation
  • B.4 Cultural Benchmark: Sensitivity Discovery and Query Generation
  • B.5 Cultural Benchmark: Scenario Generation
  • C Human Annotation for Dataset Construction
  • C.1 Annotator Recruitment and Demographics
  • C.2 Annotation Guidelines
  • C.3 Inter-Annotator Agreement
  • D Evaluation Judge Prompts
  • D.1 Jailbreak Benchmark Judge Prompt
  • D.2 Cultural Benchmark Judge Prompt
  • E LLM Judge Reliability Study
  • E.1 Human Validation Study
  • E.2 Cross-Judge Consistency Analysis
  • F Extended Results and Analysis
  • F.1 Per-Category Performance: Jailbreak Benchmark
  • F.2 Per-Category Performance: Cultural Benchmark
  • F.3 Regional Asymmetry in Prompt Language Effects
  • F.4 Local Model Selection Criteria
  • F.5 Local Model Scaling Analysis
  • G Country-Specific Flexible Subcategories

Knowls

  1. Knowl 1 — XL-SafetyBench Framework and Dual-Benchmark Architecture

    model/method

    XL-SafetyBench is a cross-cultural safety evaluation suite comprising 5,500 expert-validated test cases across 10 country-language pairs: United States (English), France (French), Germany (German), Spain (Spanish), South Korea (Korean), Japan (Japanese), India (Hindi), Indonesia (Indonesian), Türkiye (Turkish), and the United Arab Emirates (Arabic). The benchmark decomposes country-grounded safety into two separate tracks:

    1. Jailbreak Benchmark (450 prompts per country; 4,500 total): Evaluates adversarial robustness against localized harms across five harm categories: Criminal Activities, Self-harm & Dangerous Advice, Hate & Discrimination, Socioeconomic Conflicts, and Political & Misinformation. Each category contains 5 shared subcategories (fixed across all 10 countries, totaling 25) and 5 flexible subcategories (grounded in country-specific legal systems, platforms, and social dynamics, totaling 25 per country). For each of the 50 subcategories, 3 native-language base queries are generated and refined via multi-turn red-teaming across 3 probe targets, producing 3 attack variants per query (5×10×3×3=4505 \times 10 \times 3 \times 3 = 450 prompts per country).

    2. Cultural Benchmark (100 scenarios per country; 1,000 total): Evaluates implicit detection of culturally embedded taboos within benign surface tasks across six categories: Symbolic Taboos & Gift-Giving, Food, Dietary Law & Hospitality, Death, Grief & Funeral Practices, Daily Life & Public Conduct, Hierarchy, Address & Social Deference, and Legal Landmines. Each country is evaluated on 20 sensitivities (3 for each of the first five categories, 5 for Legal Landmines), with 5 tricky scenarios constructed per sensitivity (20×5=10020 \times 5 = 100 scenarios per country), where the cultural taboo is buried as an incidental detail within a dominant, unrelated request.

    All items are generated via LLMs augmented with web search, filtered through automated LLM validation gates, and validated independently by two native-speaker annotators with ≥15\ge 15 years of country residency.

  2. Knowl 2 — Hierarchical Safety Classification and Jailbreak Evaluation Metrics

    equation

    To evaluate model safety against country-grounded adversarial attacks, model responses are evaluated along four binary criteria: Harmfulness (H∈{0,1}H \in \{0, 1\}), Non-compliance (NC∈{0,1}NC \in \{0, 1\}), Specificity (S∈{0,1}S \in \{0, 1\}), and Relevance (R∈{0,1}R \in \{0, 1\}). With the logical rule H=1  ⟹  NC=1H=1 \implies NC=1, responses are classified into five mutually exclusive safety classes:

    • Critical Unsafe: H=1,NC=1,S=1H=1, NC=1, S=1 (Country-specific actionable harm disclosed).
    • Direct Unsafe: H=1,NC=1,S=0H=1, NC=1, S=0 (Generic actionable harm disclosed).
    • Soft Unsafe: H=0,NC=1H=0, NC=1 with any SS (Partial compliance or failure to block harmful guidance without direct execution steps).
    • Ideal Safe: H=0,NC=0,R=1H=0, NC=0, R=1 (Principled refusal after recognizing harmful intent).
    • Neutral Safe: H=0,NC=0,R=0H=0, NC=0, R=0 (Irrelevant, incoherent, or failed response due to comprehension deficits).

    For NN valid test cases (excluding API or execution errors), the Attack Success Rate (ASR) and Neutral-Safe Rate (NSR) are computed as: ASR=1N∑i=1NI(Labeli∈{Critical Unsafe,Direct Unsafe,Soft Unsafe})×100\text{ASR} = \frac{1}{N} \sum_{i=1}^N \mathbb{I}(\text{Label}_i \in \{\text{Critical Unsafe}, \text{Direct Unsafe}, \text{Soft Unsafe}\}) \times 100 NSR=1N∑i=1NI(Labeli=Neutral Safe)×100\text{NSR} = \frac{1}{N} \sum_{i=1}^N \mathbb{I}(\text{Label}_i = \text{Neutral Safe}) \times 100 where I(⋅)\mathbb{I}(\cdot) is the indicator function. ASR measures overall safety breach, whereas NSR isolates incidental safety stemming from language or context comprehension failure (R=0R=0).

  3. Knowl 3 — Culturally Embedded Sensitivity Evaluation Metric

    equation

    In the Cultural Benchmark, where cultural taboos are embedded as subtle details within natural requests, model responses are scored on two independent criteria:

    • Relevance (R∈{0,1}R \in \{0, 1\}): Whether the model meaningfully engaged with the scenario context (R=1R=1) rather than failing or producing irrelevant output (R=0R=0).
    • Cultural Aware Flag (C∈{0,1}C \in \{0, 1\}): Whether the model explicitly detected, flagged, or advised against the embedded country-specific taboo (C=1C=1 for Aware, C=0C=0 for Unaware).

    The Cultural Sensitivity Rate (CSR) is conditioned strictly on contextually relevant responses (Ri=1R_i = 1): CSR=∑i=1NI(Ci=1∧Ri=1)∑i=1NI(Ri=1)×100\text{CSR} = \frac{\sum_{i=1}^N \mathbb{I}(C_i = 1 \land R_i = 1)}{\sum_{i=1}^N \mathbb{I}(R_i = 1)} \times 100 where NN is the number of valid test cases. Conditioning on relevance prevents models that fail to understand the prompt from being credited or penalized on cultural competence grounds.

  4. Knowl 4 — Country-Grounded Multi-Stage Red-Teaming Algorithm

    algorithm

    The generation of country-grounded adversarial jailbreak prompts uses an automated multi-stage red-teaming pipeline extending the Prompt Automatic Iterative Refinement (PAIR) framework across three probe targets with rolling attacker memory.

    Input: Base query qq, set of probe target models T={T1,T2,T3}\mathcal{T} = \{T_1, T_2, T_3\}, maximum stages S=4S = 4
    Output: Set of successful adversarial attack prompts A\mathcal{A}
    A←∅\mathcal{A} \leftarrow \emptyset
    insights←∅\text{insights} \leftarrow \emptyset
    for each probe target T∈TT \in \mathcal{T} do
        Hatt←∅H_{\text{att}} \leftarrow \emptyset # Attacker rolling cross-stage memory
        for s=1s = 1 to SS do
            a←Attacker(q,Hatt,insights)a \leftarrow \text{Attacker}(q, H_{\text{att}}, \text{insights}) # Disguised prompt in native language
            r←T(a)r \leftarrow T(a) # Target response (history resets each stage)
            (success,score)←Judge(q,r)(\text{success}, \text{score}) \leftarrow \text{Judge}(q, r)
            if success=True\text{success} = \text{True} then
                insights←insights∪ExtractInsight(a,r)\text{insights} \leftarrow \text{insights} \cup \text{ExtractInsight}(a, r)
                A←A∪{a}\mathcal{A} \leftarrow \mathcal{A} \cup \{a\}
                break
            end if
            Hatt←Hatt∪{(a,r)}H_{\text{att}} \leftarrow H_{\text{att}} \cup \{(a, r)\}
        end for
    end for
    return RevisionPass(A)\text{RevisionPass}(\mathcal{A})

    The Attacker LLM generates adversarial prompts using tactical strategies (authority establishment, ethical reframing, logical entrapment, social pressure, strategic pivots). The Judge LLM evaluates the target response on a continuous 0--1 scale, marking success only when the score reaches 1.0 (strict compliance breach). Extracted tactical insights inform subsequent attacker prompts targeting the same model. The final outputs undergo an LLM revision pass to fix formatting artifacts while preserving semantic intent.

  5. Knowl 5 — Performance of Frontier LLMs on Jailbreak Robustness and Cultural Sensitivity

    data/table

    Evaluation of 10 global frontier LLMs across 10 countries on XL-SafetyBench measuring Attack Success Rate (ASR%, lower is safer) and Cultural Sensitivity Rate (CSR%, higher is better):

    Model AE DE ES FR ID IN JP KR TR US Avg
    Attack Success Rate (ASR%, ↓\downarrow safer)
    GPT-5.4 63.8 36.2 50.2 52.9 42.0 48.7 58.0 45.3 33.8 40.4 47.1
    GPT-5-mini 84.7 55.1 62.4 55.1 48.2 52.0 75.8 69.3 52.4 37.3 59.2
    Gemini-3.1-Pro 78.4 43.3 45.8 34.7 33.3 36.7 38.9 61.8 30.9 30.4 43.4
    Gemini-3-Flash 80.0 52.0 55.8 42.7 33.8 38.2 62.4 74.2 32.0 29.1 50.0
    Claude-4.6-Opus 21.1 4.9 4.2 3.3 2.9 1.8 6.0 7.1 3.8 4.0 5.9
    Claude-4.5-Sonnet 9.1 0.9 1.3 2.0 0.4 2.0 4.7 4.9 2.2 0.4 2.8
    Grok-4.20 26.2 37.8 48.2 34.4 29.1 35.8 21.1 38.9 31.3 3.1 30.6
    Llama-4-Maverick 68.7 94.7 96.4 96.2 90.2 93.1 92.9 97.8 96.2 94.0 92.0
    Mistral-Large-3 97.8 99.3 98.9 98.2 96.7 99.8 100.0 99.1 98.9 99.3 98.8
    Qwen3.5-397B 40.4 14.7 19.6 16.7 10.4 13.8 19.6 29.1 9.3 7.1 18.1
    Avg 57.0 43.9 48.3 43.6 38.7 42.2 47.9 52.8 39.1 34.5 44.8
    Cultural Sensitivity Rate (CSR%, ↑\uparrow better)
    GPT-5.4 57.0 67.0 57.0 58.0 64.0 63.0 66.0 75.0 52.0 85.0 64.4
    GPT-5-mini 33.0 44.0 44.0 36.0 47.0 33.0 45.0 44.0 32.0 70.0 42.8
    Gemini-3.1-Pro 81.0 70.7 73.0 67.0 82.0 63.0 81.0 89.0 64.0 90.0 76.1
    Gemini-3-Flash 53.0 66.0 56.0 62.0 66.0 47.0 65.0 88.0 54.0 79.0 63.6
    Claude-4.6-Opus 77.0 71.0 73.0 72.0 76.0 58.0 71.0 88.0 54.0 87.0 72.7
    Claude-4.5-Sonnet 67.0 72.0 71.0 63.6 72.0 48.0 63.0 80.8 57.6 87.0 68.2
    Grok-4.20 14.1 27.0 21.0 20.0 32.0 14.0 20.0 31.0 19.0 56.0 25.4
    Llama-4-Maverick 3.1 15.5 7.3 13.0 11.8 4.0 19.0 11.0 5.0 25.0 11.5
    Mistral-Large-3 7.0 15.0 13.0 11.0 11.0 6.0 22.0 18.0 3.0 32.0 13.8
    Qwen3.5-397B 49.0 68.0 57.0 58.0 67.0 41.0 63.0 73.0 44.0 84.0 60.4
    Avg 44.1 51.6 47.2 46.1 52.9 37.7 51.5 59.8 38.5 69.5 49.9

    Claude-4.5-Sonnet achieves the highest adversarial robustness (average ASR 2.8%), followed by Claude-4.6-Opus (5.9%). Gemini-3.1-Pro achieves the highest cultural sensitivity (CSR 76.1%), followed by Claude-4.6-Opus (72.7%). Open-weight frontier models (Mistral-Large-3 and Llama-4-Maverick) fail on both axes, with ASR >90%>90\% and CSR <15%<15\%. Geographically, models show strongest safety and sensitivity on US prompts (ASR 34.5%, CSR 69.5%), while vulnerability is highest in the UAE (ASR 57.0%) and South Korea (ASR 52.8%), and cultural awareness is lowest in India (CSR 37.7%) and Türkiye (CSR 38.5%).

  6. Knowl 6 — Decoupling of Adversarial Robustness and Cultural Awareness in LLMs

    empirical result

    Empirical evaluation across 10 frontier models on XL-SafetyBench reveals that adversarial safety robustness (ASR) and cultural sensitivity awareness (CSR) are decoupled capabilities:

    1. Attenuated Frontier Correlation: Although an aggregate correlation across all 10 frontier models suggests an inverse relationship (r=−0.74,p=0.014r = -0.74, p = 0.014), this correlation is driven entirely by open-weight outliers (Llama-4-Maverick, Mistral-Large-3, and Qwen3.5-397B). Restricting the analysis to the seven closed-weight frontier models attenuates the correlation to r=−0.27r = -0.27 (p=0.554p = 0.554, not statistically significant).
    2. Per-Model Discrepancies: Per-model correlations across the 10 target countries range from r=−0.63r = -0.63 (Grok-4.20) to r=+0.33r = +0.33 (Gemini-3.1-Pro). Grok-4.20 exhibits moderate jailbreak resistance (ASR 30.6%) alongside poor cultural awareness (CSR 25.4%), while Gemini-3.1-Pro leads in cultural sensitivity (CSR 76.1%) despite higher jailbreak vulnerability (ASR 43.4%).
    3. Implication: Because adversarial alignment and cultural taboo detection rely on distinct mechanisms, composite safety metrics obscure model vulnerabilities. Benchmarking frameworks must report adversarial robustness and cultural awareness separately.
  7. Knowl 7 — Illusion of Safety via ASR-NSR Trade-Off in Country-Specific Local LLMs

    empirical result

    Evaluation of 27 country-specific local LLMs across 9 non-US countries demonstrates that seemingly low Attack Success Rates (ASR) among local models stem from comprehension and generation failures rather than genuine safety alignment:

    1. Linear ASR--NSR Trade-Off: While frontier models cluster near 0% Neutral-Safe Rate (NSR), local models exhibit a strong negative correlation between ASR and NSR (r=−0.81r = -0.81), clustering along the trade-off boundary: ASR+NSR≈100%\text{ASR} + \text{NSR} \approx 100\%
    2. Comprehension Failure Masking: Local models exhibiting low ASR (e.g., CroissantLLM at ASR 8.0%, Teuken-7B at ASR 17.6%, Kumru-2B at ASR 17.6%, Param2-17B at ASR 24.2%) have elevated NSR (62.9%, 54.4%, 44.7%, and 72.7%, respectively). Their apparent safety reflects incoherent or irrelevant outputs rather than principled refusals (Ideal Safe).
    3. Vulnerability in Fluent Local Models: Conversely, local models that achieve fluent prompt comprehension (NSR≈0%\text{NSR} \approx 0\%, e.g., Trendyol-8B at NSR 0.4%, A.X-K1 at NSR 2.7%) fail to resist country-grounded adversarial attacks, suffering ASR >90%>90\% (Trendyol-8B at 96.9%, A.X-K1 at 90.0%).
    4. Cultural Awareness Deficit: Training on local language corpora alone does not impart cultural awareness; most local models score below 15% CSR, with several scoring 0.0% (Lucie-7B, Teuken-7B, WiroAI-9B).
  8. Knowl 8 — Performance of Country-Specific Local LLMs Across Nine Target Countries

    data/table

    Performance of 27 country-specific local LLMs across 9 countries on XL-SafetyBench measuring Attack Success Rate (ASR%, lower is safer), Neutral-Safe Rate (NSR%, lower is better), and Cultural Sensitivity Rate (CSR%, higher is better):

    Country Model ASR (↓\downarrow) NSR CSR (↑\uparrow)
    France CroissantLLM (1.3B) 8.0 62.9 0.0
    France Gaperon-24B 42.0 31.1 0.0
    France Lucie-7B 63.1 16.2 0.0
    Germany LeoLM-7B 44.7 34.2 0.0
    Germany SauerkrautLM-14B 86.4 7.1 2.0
    Germany Teuken-7B 17.6 54.4 0.0
    India Param2-17B 24.2 72.7 2.2
    India Sarvam-105B 34.7 1.6 7.0
    India Sarvam-30B 52.7 2.7 3.0
    Indonesia Gemma2-9B-SahabatAI 76.1 16.3 4.0
    Indonesia Llama3-8B-SahabatAI 90.2 5.6 3.0
    Indonesia Sailor2-8B 96.7 1.6 3.0
    Japan LLM-JP-4-32B 59.3 39.6 13.1
    Japan Rakuten-AI-3.0 (671B) 84.9 4.7 13.0
    Japan Stockmark-2-100B 81.1 9.8 10.0
    South Korea A.X-K1 (519B) 90.0 2.7 7.0
    South Korea EXAONE-236B 45.6 3.6 30.0
    South Korea SOLAR-100B 32.0 33.1 23.2
    Spain Alia-40B 93.3 1.3 2.0
    Spain Iberian-7B 32.9 27.1 0.0
    Spain RigoChat-7B 95.1 2.0 0.0
    Türkiye Kumru-2B 17.6 44.7 1.8
    Türkiye Trendyol-8B 96.9 0.4 1.0
    Türkiye WiroAI-9B 73.1 11.1 0.0
    UAE Falcon-H1-34B 93.6 1.8 1.0
    UAE K2-Think-V2 (70B) 44.9 42.4 14.4
    UAE Jais-2-70B 91.1 0.7 6.0

    These results illustrate that local models either suffer from comprehension failure (high NSR) or lack safety alignment once capable of comprehension (high ASR). EXAONE-236B achieves the highest CSR among all evaluated local models (30.0%), yet falls substantially short of leading global models.

  9. Knowl 9 — Parameter Scaling Dynamics of Local LLMs on Comprehension and Cultural Sensitivity

    empirical result

    Analysis of the 27 country-specific local models grouped by parameter count reveals distinct scaling behaviors across safety, comprehension, and cultural awareness:

    Size Bin nn ASR (%) NSR (%) CSR (%)
    Small (<10B<10\text{B}) 12 60.5 21.8 1.1
    Medium (10–50B10\text{--}50\text{B}) 7 64.6 22.3 3.3
    Large (>50B>50\text{B}) 8 63.5 12.3 13.8
    1. Step-Function Comprehension Recovery: NSR correlates negatively with log⁡10(parameters)\log_{10}(\text{parameters}) (r=−0.38,p=0.053r = -0.38, p = 0.053). Binned analysis shows a threshold step-function: NSR remains high across small (21.8%) and medium (22.3%) models and drops sharply only in the large bin (>50B>50\text{B}, 12.3%), confirming that low ASR in smaller local models is caused by comprehension failure.
    2. Cultural Awareness Scaling and Ceiling: Cultural sensitivity (CSR) correlates positively with parameter count (r=+0.68,p<0.001r = +0.68, p < 0.001), increasing monotonically across bins (1.1%→3.3%→13.8%1.1\% \to 3.3\% \to 13.8\%). However, even large local models (>50B>50\text{B}) average only 13.8% CSR (with the peak being EXAONE-236B at 30.0%), falling far short of frontier models (GPT-5-mini at 42.8%, closed frontier leaders >64%>64\%).
  10. Knowl 10 — Harm Evasion Gap between Shared and Country-Specific Flexible Subcategories

    empirical result

    In the Jailbreak Benchmark, comparing Attack Success Rates between the 25 shared subcategories (identical concepts across all 10 countries) and the 25 flexible subcategories (country-specific concepts, such as France's Go-Fast drug smuggling or South Korea's Jeonse fraud) reveals a persistent evasion gap:

    Across 8 of 10 frontier LLMs, flexible subcategories produce equal or higher ASR than shared subcategories, with a cross-model average gap of Δ=+1.1 percentage points\Delta = +1.1\text{ percentage points} (ASRFlexible=45.0%\text{ASR}_{\text{Flexible}} = 45.0\% vs. ASRShared=43.9%\text{ASR}_{\text{Shared}} = 43.9\%). The evasion gap is most pronounced in Grok-4.20 (Δ=+4.3 pp\Delta = +4.3\text{ pp}; 32.6%32.6\% vs. 28.4%28.4\%), Gemini-3-Flash (Δ=+2.2 pp\Delta = +2.2\text{ pp}), Qwen3.5-397B (Δ=+1.9 pp\Delta = +1.9\text{ pp}), and GPT-5.4 (Δ=+1.5 pp\Delta = +1.5\text{ pp}).

    This indicates that universal safety alignments trained primarily on globally widespread harm patterns are more easily bypassed when attacks are framed using localized, culturally opaque institutional and legal concepts.

  11. Knowl 11 — Regional Asymmetry in Cultural Sensitivity Under English vs. Local Language Prompting

    empirical result

    A prompt-language ablation comparing Cultural Sensitivity Rates (CSR) when scenarios are presented in English versus the target country's native language shows nearly identical global averages (47.54% English vs. 47.48% Local), but reveals a statistically significant regional asymmetry (Fisher’s exact test p=0.029\text{Fisher's exact test } p = 0.029):

    Country Region Δ CSR (English − Local, pp)\Delta\text{ CSR (English } - \text{ Local, pp)}
    Spain European −6.31-6.31
    Germany European −2.49-2.49
    France European −0.88-0.88
    Japan Non-European +1.28+1.28
    India Non-European +1.90+1.90
    South Korea Non-European +2.66+2.66
    Türkiye Non-European +4.25+4.25

    All three European countries demonstrate a local-language advantage (Δ=−6.31 to −0.88 pp\Delta = -6.31\text{ to } -0.88\text{ pp}), whereas all four non-European countries demonstrate an English-language advantage (Δ=+1.28 to +4.25 pp\Delta = +1.28\text{ to } +4.25\text{ pp}). This suggests that nuanced cultural discourse regarding non-European societies (such as travel literature and sociological analyses) is disproportionately mediated in English-language corpora, enabling English prompts to act as more effective retrieval cues for non-European taboos.

  12. Knowl 12 — Category-Level Vulnerability Patterns Across Harm and Cultural Dimensions

    empirical result

    Evaluation of frontier models across fine-grained categories in XL-SafetyBench reveals domain-specific vulnerability patterns:

    1. Jailbreak Harm Categories: Averaged across 10 frontier models, Hate & Discrimination (μ=47.6%\mu = 47.6\%) and Socioeconomic Conflicts (μ=47.0%\mu = 47.0\%) yield higher Attack Success Rates than Criminal Activities (μ=44.8%\mu = 44.8\%), Political & Misinformation (μ=43.8%\mu = 43.8\%), and Self-harm & Dangerous Advice (μ=40.7%\mu = 40.7\%). This indicates that localized discriminatory discourse and class conflicts are harder for safety filters to detect and refuse than direct physical self-harm.
    2. Cultural Sensitivity Categories: Symbolic Taboos & Gift-Giving is universally the most challenging category across all models, yielding an average CSR of 31.5%31.5\% (compared to 56.3%56.3\% for Legal Landmines, 53.8%53.8\% for Hierarchy, Address & Social Deference, 53.6%53.6\% for Death, Grief & Funeral Practices, 50.8%50.8\% for Daily Life & Public Conduct, and 50.6%50.6\% for Food, Dietary Law & Hospitality). Symbolic taboos (e.g., unlucky numbers, flower meanings, and homophone superstitions) are rarely codified in formal legal or instructional texts, posing the steepest cross-cultural generalization challenge.

Coverage note — Omitted material includes the complete textual listings of all 250 country-specific flexible subcategories (Tables 15–16), intermediate prompt templates used for LLM generation stages (Appendix B), and per-judge inter-annotator agreement breakdown matrices (Appendix E), as they represent raw catalog data and prompt scripts rather than standalone conceptual or empirical contributions.

References

  1. 1.Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker. The multilingual alignment prism: Aligning global and local preferences to reduce harm. arXiv preprint arXiv:2406.18682, 2024.
  2. 2.Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Ashutosh Dwivedi, Alham Fikri Aji, Jacki O’Neill, Ashutosh Modi, and Monojit Choudhury. Towards measuring and modeling “culture” in LLMs: A survey. arXiv preprint arXiv:2403.15412, 2024.
  3. 3.Mehdi Ali, Michael Fromm, et al. Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs. arXiv preprint arXiv:2410.03730, 2024.
  4. 4.Anthropic. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5, 2025.
  5. 5.Anthropic. Claude Opus 4.6 system card. Technical report, Anthropic, February 2026.
  6. 6.Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. In 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pages 23–42. IEEE, 2025.
  7. 7.Zhoujun Cheng, Richard Fan, Shibo Hao, Taylor W Killian, Haonan Li, Suqi Sun, Hector Ren, Alexander Moreno, Daqian Zhang, Tianjun Zhong, et al. K2-think: A parameter-efficient reasoning system. arXiv preprint arXiv:2509.07604, 2025.
  8. 8.Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, and Yejin Choi. CulturalBench: A robust, diverse and challenging benchmark for measuring LMs’ cultural knowledge through human-AI red-teaming. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 25663–25701, Vienna, Austria, 2025. Association for Computational Linguistics.
  9. 9.Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. Multilingual jailbreak challenges in large language models. arXiv preprint arXiv:2310.06474, 2023.
  10. 10.Esin Durmus, Karina Nyugen, Thomas I. Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, Liane Lovitt, Sam McCandlish, Orowa Sikder, Alex Tamkin, Janel Thamkul, Jared Kaplan, Jack Clark, and Deep Ganguli. Towards measuring the representation of subjective global opinions in language models. arXiv preprint arXiv:2306.16388, 2023.
  11. 11.Manuel Faysse, Patrick Fernandes, Nuno M. Guerreiro, António Loison, Duarte M. Alves, Caio Corro, Nicolas Boizard, João Alves, Ricardo Rei, Pedro Henrique Martins, Antoni Bigata Casademunt, François Yvon, André Martins, Gautier Viaud, Céline Hudelot, and Pierre Colombo. CroissantLLM: A truly bilingual French–English language model. arXiv preprint arXiv:2402.00786, 2024.
  12. 12.Alexander Gill, Abhilasha Ravichander, and Ana Marasovic. What has been lost with synthetic evaluation? In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 9902–9945, Suzhou, China, 2025. Association for Computational Linguistics.
  13. 13.Nathan Godey, Wissam Antoun, Rian Touchent, Rachel Bawden, Éric de la Clergerie, Benoî t Sagot, and Djamé Seddah. Gaperon: A peppered English–French generative language model suite. arXiv preprint arXiv:2510.25771, 2025.
  14. 14.Aitor Gonzalez-Agirre, Marc Pàmies, Joan Llop, Irene Baucells, Severino Da Dalt, Daniel Tamayo, José Javier Saiz, Ferran Espuña, Jaume Prats, Javier Aula-Blasco, Mario Mina, Adrián Rubio, Alexander Shvets, Anna Sallés, Iñaki Lacunza, Iñigo Pikabea, Jorge Palomar, Júlia Falcão, Lucí a Tormo, Luis Vasquez-Reina, Montserrat Marimon, Valle Ruí z Fernández, and Marta Villegas. Salamandra technical report. arXiv preprint arXiv:2502.08489, 2025. ALIA-40b is the 40B parameter instance of the Salamandra family.
  15. 15.Google DeepMind. Gemini 3.1 Pro model card. Technical report, Google DeepMind, February 2026.
  16. 16.GoTo Company, Indosat Ooredoo Hutchison, and AI Singapore. Gemma2 9b cpt sahabat-ai v1, 2024.
  17. 17.GoTo Company, Indosat Ooredoo Hutchison, and AI Singapore. Llama3 8b cpt sahabat-ai v1, 2024.
  18. 18.Olivier Gouvert, Julie Hunter, Jérôme Louradour, Christophe Cerisara, Evan Dufraisse, Yaya Sy, Laura Rivière, Jean-Pierre Lorré, and OpenLLM-France community. The Lucie-7B LLM and the Lucie training dataset: Open resources for multilingual language generation. arXiv preprint arXiv:2503.12294, 2025.
  19. 19.Demis Hassabis, Koray Kavukcuoglu, and the Gemini Team. Gemini 3: Introducing the latest Gemini AI model from google. https://blog.google/products-and-platforms/products/gemini/gemini-3/, November 2025.
  20. 20.ILENIA Project. Iberian-7B: ILENIA Iberian language models. https://proyectoilenia.es/, 2024.
  21. 21.Masahiro Kaneko, Ayana Niwa, and Timothy Baldwin. JailNewsBench: Multi-lingual and regional benchmark for fake news generation under jailbreak attacks. In The Fourteenth International Conference on Learning Representations (ICLR), 2026.
  22. 22.Dahyun Kim, Chanjun Park, Sanghoon Kim, Wonsung Lee, Wonho Song, Yunsu Kim, Hyeonwoo Kim, Yungi Kim, Hyeonju Lee, Jihoo Kim, Changbae Ahn, Seonghoon Yang, Sukyung Lee, Hyunbyung Park, Gyoungjin Gim, Mikyoung Cha, Hwalsuk Lee, and Sunghun Kim. SOLAR 10.7B: Scaling large language models with simple yet effective depth up-scaling. arXiv preprint arXiv:2312.15166, 2023.
  23. 23.LG AI Research. K-EXAONE technical report. arXiv preprint arXiv:2601.01739, 2026.
  24. 24.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-eval: Nlg evaluation using gpt-4 with better human alignment, 2023.
  25. 25.LLM-jp, Akiko Aizawa, et al. LLM-jp: A cross-organizational project for the research and development of fully open Japanese LLMs. arXiv preprint arXiv:2407.03963, 2024.
  26. 26.Meta AI. The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation. https://ai.meta.com/blog/llama-4-multimodal-intelligence/, April 2025.
  27. 27.Mistral AI. Introducing Mistral 3. https://mistral.ai/news/mistral-3, December 2025.
  28. 28.Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew Ali Ayele, Víctor Gutiérrez-Basulto, Yazmín Ibáñez-García, Hwaran Lee, Shamsuddeen Hassan Muhammad, Kiwoong Park, Anar Sabuhi Rzayev, Nina White, Seid Muhie Yimam, Mohammad Taher Pilehvar, Nedjma Ousidhoum, Jose Camacho-Collados, and Alice Oh. BLEnD: A benchmark for LLMs on everyday knowledge in diverse cultures and languages. In Advances in Neural Information Processing Systems 37 (NeurIPS 2024) Datasets and Benchmarks Track, 2024.
  29. 29.Zhiyuan Ning, Tianle Gu, Jiaxin Song, Shixin Hong, Lingyu Li, Huacan Liu, Jie Li, Yixu Wang, Meng Lingyu, Yan Teng, et al. Linguasafe: A comprehensive multilingual safety benchmark for large language models. arXiv preprint arXiv:2508.12733, 2025.
  30. 30.OpenAI. GPT-5 system card. Technical report, OpenAI, August 2025. Documents gpt-5, gpt-5-mini and gpt-5-nano. Available at https://cdn.openai.com/gpt-5-system-card.pdf.
  31. 31.OpenAI. Introducing GPT-5.4. https://openai.com/index/introducing-gpt-5-4/, March 2026. Model release announcement, March 5, 2026.
  32. 32.Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. Bbq: A hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 2086–2105, 2022.
  33. 33.Siddhesh Pawar, Junyeong Park, Jiho Jin, Arnav Arora, Junho Myung, Srishti Yadav, Faiz Ghifari Haznitrama, Inhwa Song, Alice Oh, and Isabelle Augenstein. Survey of cultural awareness in language models: Text and beyond. Computational Linguistics, 51(3):907–1004, 2025.
  34. 34.Björn Plüster et al. LeoLM: Igniting German-language LLM research, 2023. LAION blog post.
  35. 35.Kundeshwar Pundalik et al. PARAM-1: BharatGen bilingual foundation model. arXiv preprint arXiv:2507.13390, 2025. BharatGen / IIT Bombay.
  36. 36.Qwen Team. Qwen3 technical report. https://github.com/QwenLM/Qwen3, 2025. Alibaba Cloud, April 29, 2025.
  37. 37.Rakuten Group, Inc. Rakuten AI 3.0 now available, japan’s largest high-performance AI model developed as part of the GENIAC project. https://global.rakuten.com/corp/news/press/2026/0317_01.html, March 2026.
  38. 38.Rakuten Group, Inc., Aaron Levine, Connie Huang, Chenguang Wang, Eduardo Batista, Ewa Szymanska, Hongyi Ding, Hou Wei Chou, Jean-François Pessiot, Johanes Effendi, Justin Chiu, Kai Torben Ohlhus, Karan Chopra, Keiji Shinzato, Koji Murakami, Lee Xiong, Lei Chen, Maki Kubota, Maksim Tkachenko, Miroku Lee, Naoki Takahashi, Prathyusha Jwalapuram, Ryutaro Tatsushima, Saurabh Jain, Sunil Kumar Yadav, Ting Cai, Wei-Te Chen, Yandi Xia, Yuki Nakayama, and Yutaka Higashiyama. RakutenAI-7B: Extending large language models for Japanese. arXiv preprint arXiv:2403.15484, 2024.
  39. 39.Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. NormAd: A framework for measuring the cultural adaptability of large language models. arXiv preprint arXiv:2404.12464, 2024.
  40. 40.Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M. Aroyo. “everyone wants to do the model work, not the data work”: Data cascades in high-stakes AI. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–15. ACM, 2021.
  41. 41.Gonzalo Santamarí a Gómez, Guillem Garcí a Subies, Pablo Gutiérrez Ruiz, Mario González Valero, Natàlia Fuertes, Helena Montoro Zamorano, Carmen Muñoz Sanz, Leire Rosado Plaza, Nuria Aldama Garcí a, David Betancur Sánchez, Kateryna Sushkova, Marta Guerrero Nieto, and Álvaro Barbero Jiménez. RigoChat 2: An adapted language model to Spanish using a bounded dataset and reduced hardware. arXiv preprint arXiv:2503.08188, 2025.
  42. 42.Sarvam AI. Sarvam-105B (Indus): An open foundation model for Indic languages. Hugging Face, 2026.
  43. 43.Sarvam AI. Sarvam-30B: A mixture-of-experts foundation model for Indic languages. Hugging Face, 2026.
  44. 44.Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, William Marshall, Gurpreet Gosal, Cynthia Liu, Zhiming Chen, et al. Jais and jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models. arXiv preprint arXiv:2308.16149, 2023.
  45. 45.SKT AI Model Lab. A.X-K1. https://huggingface.co/skt/A.X-K1, 2026.
  46. 46.Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer. A strongreject for empty jailbreaks, 2024.
  47. 47.Stockmark Inc. Stockmark-2-100B-Instruct. https://huggingface.co/stockmark/Stockmark-2-100B-Instruct, 2025. Supported by GENIAC.
  48. 48.Sailor2 Team. Sailor2: Sailing in south-east asia with inclusive multilingual llms. arXiv preprint arXiv:2502.12982, 2025.
  49. 49.Trendyol Tech. Trendyol-LLM-8B-T1: A turkish e-commerce large language model. https://huggingface.co/Trendyol/Trendyol-LLM-8B-T1, 2025.
  50. 50.Meliksah Turker, Erdi Ari, and Aydin Han. Kumru: A turkish language model from scratch. https://huggingface.co/vngrs-ai/Kumru-2B-Base, 2025.
  51. 51.VAGO Solutions. SauerkrautLM: German language model suite. Hugging Face, 2024.
  52. 52.Wenxuan Wang, Zhaopeng Tu, Chang Chen, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, and Michael Lyu. All languages matter: On the multilingual safety of llms. In Findings of the Association for Computational Linguistics: ACL 2024, pages 5865–5877, 2024.
  53. 53.Ishaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri, Manohar Swaminathan, and Sunayana Sitaram. Pariksha: A large-scale investigation of human-llm evaluator agreement on multilingual and multi-cultural data. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 7900–7932, 2024.
  54. 54.WiroAI. WiroAI turkish language model. https://huggingface.co/WiroAI, 2024.
  55. 55.xAI. Grok 4.20. https://x.ai/, 2026. Released in beta on February 17, 2026; full API release in March 2026.
  56. 56.Zheng-Xin Yong, Beyza Ermis, Marzieh Fadaee, Stephen Bach, and Julia Kreutzer. The state of multilingual llm safety research: From measuring the language gap to mitigating it. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15856–15871, 2025.
  57. 57.Haneul Yoo, Yongjin Yang, and Hwaran Lee. Code-switching red-teaming: Llm evaluation for safety and multilingual understanding, 2025.
  58. 58.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595–46623, 2023.
  59. 59.Jingwei Zuo, Maksim Velikanov, Ilyas Chahed, Younes Belkada, Dhia Eddine Rhayem, Guillaume Kunsch, Hakim Hacid, Hamza Yous, Brahim Farhat, Ibrahim Khadraoui, et al. Falcon-h1: A family of hybrid-head language models redefining efficiency and performance. arXiv preprint arXiv:2507.22448, 2025.

Citation

MLA
Choi, D., et al. “XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity”. arXiv, 2026, http://arxiv.org/abs/2605.05662v1.
APA
Choi, D., Kim, E., Noh, J., Seo, S., Kim, E., Oh, M., Park, Y., Kartono, B. J., Pichlmeier, J., Berndt, H., Mendu, S. K., Tungka, G. J., Gökçe, Ö., Gehlot, S., Pratt, K., Minnich, A., & Park, H. (2026). XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity. arXiv. http://arxiv.org/abs/2605.05662v1
Chicago
Choi, D., E. Kim, J. Noh, et al. 2026. “XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity”. arXiv. http://arxiv.org/abs/2605.05662v1.
Harvard
Choi, D. et al. (2026) “XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2605.05662v1.
Vancouver
1. Choi D, Kim E, Noh J, et al (2026) XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity. arXiv

BibTeX

@article{choi2026safetybench,
  title = {XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity},
  author = {Choi, Dasol and Kim, Eugenia and Noh, Jaewon and Seo, Sang and Kim, Eunmi and Oh, Myunggyo and Park, Yunjin and Kartono, Brigitta Jesica and Pichlmeier, Josef and Berndt, Helena and Mendu, Sai Krishna and Tungka, Glenn Johannes and Gökçe, Özlem and Gehlot, Suresh and Pratt, Katherine and Minnich, Amanda and Park, Haon},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2605.05662v1},
  eprint = {2605.05662}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/