TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

Victor AkinodeSenyu LiWassim HamidoucheWaqas ZamirInbal Becker-ReshefDavid Ifeoluwa Adelani

article2026arXiv3 citationsHonorable mention for best paper

Introduces TukaBench to evaluate large language model safety across seven African languages, demonstrating that cultural context and code-switching significantly reduce model refusals while exposing critical failures in automated safety judging.

Listen

Safety evaluations for large language models remain heavily centered on English and a small group of high-resource languages. This leaves low-resource languages, especially African languages spoken by tens of millions of people, critically vulnerable to adversarial manipulation and jailbreaking attacks. Standard safety benchmarks typically rely on direct translations of Western-centric scenarios, failing to capture local cultural nuances and regional harm patterns.

The article evaluates model safety in low-resource environments by introducing TukaBench, the first culturally grounded jailbreak benchmark spanning English and seven African languages: Amharic, Hausa, Igbo, Chichewa, Kiswahili, Yorùbá, and isiXhosa. The benchmark tests how language translation, local cultural context, code-switching with English, and adversarial perturbation affect safety guardrails across thirteen leading proprietary and open-weight language models.

The benchmark encompasses 986 prompts per language across five datasets, constructed via machine translation followed by native-speaker post-editing and cultural adaptation. Prompts span ten misuse categories and incorporate culturally specific entities and harms, such as local financial fraud and governance issues. To evaluate responses, the authors introduce a three-part classification scheme: Refused, Jailbroken, and Deflected. The deflection metric identifies cases where a model neither complies with nor explicitly refuses a harmful prompt, but instead fails to comprehend the input and responds off-target. Automated evaluations using an automated model judge were validated against human majority labels across 1,500 model responses.

The evaluation reveals several critical findings regarding model safety in under-resourced languages. First, prompting in African languages reduces explicit model refusals from an average of 78.1% in English to 62.7%, while deflection jumps from 6.1% to 20.1%. Second, cultural grounding significantly increases model vulnerability: culturally adapted prompts generate higher failure rates than direct translations, proving that direct translation benchmarks underestimate real-world deployment risks. Third, language resource availability and script type strongly dictate comprehension; higher-resource languages like Swahili demonstrate stronger refusal behavior, whereas lower-resource languages and non-Latin scripts like Amharic suffer from severe comprehension breakdowns. Fourth, automated evaluation reliability drops sharply in lower-resource settings, where agreement between automated judges and native human annotators falls from roughly 80% in Swahili to below 60% in Yorùbá, Igbo, and Amharic. Finally, newer model generations and adversarial attacks like boundary point jailbreaking reduce off-target deflections, but often convert these newly understood prompts into harmful completions.

These findings indicate that relying strictly on attack success rates produces a false sense of security. Low attack rates in African languages often reflect poor language comprehension rather than robust safety alignment. Furthermore, automated safety evaluation pipelines become unreliable in low-resource and non-Latin script contexts, compounding deployment risks in multilingual regions. As models improve at language comprehension, their vulnerability to harmful compliance increases unless safety alignment is trained alongside linguistic capabilities.

Organizations developing or deploying models in multilingual environments must avoid using simple direct translation for safety auditing and instead adopt culturally localized test suites. Evaluation frameworks should incorporate deflection metrics to decouple comprehension errors from genuine refusals, and automated judging should be paired with human native-speaker spot checks. Because the benchmark focuses on relatively resourced African languages and snapshot model versions, future work must expand coverage to lower-resource dialects and build dedicated multilingual safety guardrails.

Cover for TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

Abstract

Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, critically underexplored. We introduce TUKABENCH, a jailbreak benchmark for seven African languages that extends JailbreakBench (JBB) beyond direct translation through four settings: human translation of JBB prompts, English adaptation to African contexts followed by human translation, human-curated prompts validated through interactions with GPT-5.2, and code-switched prompts combining English and African languages, isolating the effect of language, cultural grounding, and prompt evasiveness on model safety. Across closed and open models, prompting in African languages reduces refusal relative to English, with culturally adapted prompts leading to least refusal. The evaluation also surfaces two structural limitations: model comprehension failures and reduced LLM-as-a-judge reliability in LRLs. To capture the first, we introduce Deflection alongside Refused and Jailbroken; to assess the second, we validate outputs with human annotations, showing that judge-human agreement drops in lower-resource languages and less commonly supported scripts.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 TukaBench
  • 3.1 Language Selection
  • 3.2 Benchmark Components
  • 3.2.1 JBB-Derived Components
  • 3.2.2 African-Authored Components
  • 3.3 Translation and Adaptation Pipeline
  • 4 Experimental Setup
  • 4.1 Models
  • 4.2 Prompting Methods
  • 4.3 LLM-as-a-Judge Evaluation
  • 4.4 Human Verification
  • 4.5 Metrics
  • 5 Results
  • 5.1 Direct Prompting Results
  • 5.2 BPJ and Code-Switched Results
  • 5.3 Human Verification and Judge Reliability
  • 5.4 Benign Prompt Evaluation
  • 6 Conclusion
  • References
  • A Benchmark Construction and Annotation
  • A.1 Annotator Recruitment and Compensation
  • A.2 Task Allocation Across the Two Annotators
  • A.3 Annotation Workflow
  • A.4 Quality Control
  • A.5 Annotator Communication
  • A.6 Annotator Statistics
  • A.7 Human Evaluation of Model Outputs
  • A.8 Code-Switched Prompt Annotation
  • B Prompt Categorization
  • B.1 Category Definitions
  • B.2 Annotation Procedure
  • B.3 Category Distribution in AfriJail
  • C Boundary Point Jailbreaking: Technical Details
  • D Prompts and Decoding Settings
  • D.1 Judge Prompt
  • D.2 Direct Prompting
  • D.3 Code-Switching Attack at Evaluation Time
  • D.4 Code-Switched Data Generation
  • E Additional Results

Knowls

  1. Knowl 1 — TukaBench Benchmark Dataset Composition and Scope

    definition

    TukaBench is a culturally grounded jailbreak evaluation benchmark consisting of 986 prompts per language across English and seven African languages representing two major African language families:

    • Afro-Asiatic: Amharic (amh, written in Ge'ez script) and Hausa (hau, written in Latin script).
    • Niger-Congo: Igbo (ibo), Chichewa/Nyanja (nya), Kiswahili (swh), Yorùbá (yor), and isiXhosa (xho), all written in Latin script.

    The benchmark organizes prompts along two design axes: source origin (derived from JailbreakBench [JBB] versus newly African-authored) and construction strategy (direct translation, cultural adaptation, or code-switching), yielding five distinct components:

    1. Afri-JBB-Harm (100 prompts): Direct human translations of the 100 harmful prompts from JailbreakBench, preserving original Western named entities and context.
    2. Afri-JBB-Benign (100 prompts): Direct human translations of the 100 benign control prompts from JailbreakBench, preserving Western context to measure over-refusal and baseline comprehension.
    3. Afri-JBB-Culture (100 prompts): JailbreakBench harmful prompts culturally adapted to African contexts by substituting Western entities with African counterparts (such as national examination boards or regional bodies), followed by human translation and localization.
    4. AfriJail-Mono (343 prompts): Culturally grounded adversarial prompts authored in English by human data collectors in Nigeria covering regional dynamics, local governance, and specific African social harms (such as mob justice, examination fraud under WAEC/JAMB, and mobile-money scams). Prompts trivially refused by GPT-5.2 were filtered out, and the remaining prompts were translated and culturally localized across all seven African languages under the ten standard JailbreakBench misuse categories.
    5. AfriJail-CS (343 prompts): Intra-sentential code-switched versions of AfriJail-Mono combining English and the target African language, constructed for six African languages (excluding isiXhosa).
  2. Knowl 2 — Three-Way Safety Evaluation Scheme for Low-Resource Languages

    definition

    In low-resource language safety evaluation, standard binary evaluation (Jailbroken vs. Refused) conflates genuine safety alignment refusals with comprehension failures where models produce incoherent or off-target outputs. TukaBench defines a three-way response classification scheme:

    • Jailbroken: The model response substantively fulfills and complies with the harmful intent of the adversarial prompt.
    • Refused: The model response constitutes an explicit refusal to comply with the harmful request (e.g., standard refusal statements such as "I cannot fulfill this request").
    • Deflected: The model response is neither an explicit refusal nor a substantive compliance, exhibiting off-target generation, incomplete answers, incoherence, or apparent failure to comprehend the input prompt.

    Over an evaluation dataset of harmful prompts, three mutually exclusive metrics are computed:

    • Attack Success Rate (ASR): The proportion of model responses classified as Jailbroken.
    • Deflection Rate: The proportion of model responses classified as Deflected.
    • Refusal Rate: The proportion of model responses classified as Refused.

    By construction, these three metrics sum to 100%100\%: ASR+Deflection Rate+Refusal Rate=100%\text{ASR} + \text{Deflection Rate} + \text{Refusal Rate} = 100\%

  3. Knowl 3 — TukaBench Translation, Cultural Localization, and Code-Switching Pipeline

    model/method

    To generate natural and culturally aligned prompts in African languages without standard machine translation artifacts, TukaBench utilizes a four-stage translation and adaptation pipeline:

    1. Machine Translation: Source English prompts are translated using Google Translate for Amharic, Hausa, Igbo, Chichewa, Kiswahili, and isiXhosa. For Yorùbá, where standard translation engines omit tonal diacritics and proprietary LLMs exhibit high refusal rates on sensitive prompts, translations are generated by AfriqueQwen-8B prompted with five few-shot examples from the MAFAND dataset under greedy decoding (temperature=0.0\text{temperature} = 0.0).
    2. Quality Estimation (QE) Pre-Screening: Machine-translated outputs are evaluated reference-free using SSA-COMET-QE, a quality estimation metric adapted from COMET for African languages. Outputs scoring below 0.500.50 (flagged as LOW) are re-translated via fallback systems prior to human review, while scores 0.50–0.640.50\text{--}0.64 are flagged as MEDIUM and ≥0.65\ge 0.65 as HIGH.
    3. Human Post-Editing and Cultural Localization: Two native speakers per language (14 annotators total, all holding at least a bachelor's degree) correct translation errors. For culturally grounded components (Afri-JBB-Culture and AfriJail-Mono), annotators localize Western named entities, institutions, currencies, and geographic references to match regional contexts (e.g., substituting Kenyan institutions for Swahili, Nigerian for Yorùbá/Hausa/Igbo). Localization is audited via English back-translations.
    4. Code-Switched Prompt Generation (AfriJail-CS): AfriqueQwen-8B generates code-switched candidates using three native-speaker few-shot exemplars (English prompt, target-language translation, and natural code-switched variant) via greedy decoding. Fluent bilingual native speakers then review and correct every generated prompt to ensure natural conversational mixing.
  4. Knowl 4 — Deflection Surge vs. Attack Success Rate in African-Language Jailbreak Evaluation

    empirical result

    Evaluating thirteen proprietary and open-weight LLMs (including GPT-3.5-Turbo, GPT-4o, GPT-5.2, Grok-3, Grok-4, Grok-4.3, Gemini-2.5-Pro, Gemini-3.1-Pro, Claude-Haiku-4.5, Claude-Sonnet-5, Claude-Opus-4.8, Gemma-3-27B, Gemma-4-31B, Qwen-3.5-27B, GPT-OSS-120B, Llama-4-Maverick, and DeepSeek-3.2) across combined harmful subsets (Afri-JBB-Harm, Afri-JBB-Culture, AfriJail-Mono) under Direct Prompting demonstrates that switching from English to African languages does not lead to a consistent increase in Attack Success Rate (ASR). Instead, the dominant cross-lingual shift is a marked surge in Deflection Rate and a corresponding drop in Refusal Rate.

    Across all evaluated models, the cross-model averages shift as follows:

    • English: ASR=15.8%\text{ASR} = 15.8\%, Refusal Rate=78.1%\text{Refusal Rate} = 78.1\%, Deflection Rate=6.1%\text{Deflection Rate} = 6.1\%.
    • African Languages (mean across six languages, excluding isiXhosa): ASR=17.2%\text{ASR} = 17.2\%, Refusal Rate=62.7%\text{Refusal Rate} = 62.7\%, Deflection Rate=20.1%\text{Deflection Rate} = 20.1\%.

    Proprietary models maintain more stable refusal and comprehension profiles, while open-weight models experience substantially higher rates of comprehension failure (e.g., DeepSeek-3.2 African-language Deflection is 20.4%20.4\%; GPT-OSS-120B is 20.9%20.9\%). Low or stable ASR in low-resource African languages reflects off-target generation and comprehension breakdown rather than effective safety alignment.

  5. Knowl 5 — Cultural Grounding vs. Direct Translation in Model Jailbreak Vulnerability

    empirical result

    Comparing direct English translations (Afri-JBB-Harm) against culturally grounded prompts (Afri-JBB-Culture and AfriJail-Mono) demonstrates that direct translation substantially underestimates real-world safety vulnerability in non-Western contexts. Culturally grounded prompts incorporate local institutions, regional conflict scenarios, and culturally specific harm patterns (such as examination cheating schemes or local financial scams), which bypass alignment filters calibrated primarily on Western concepts.

    Under Direct Prompting evaluated via GPT-4.1:

    • For GPT-4o, average African-language ASR rises from 13.5%13.5\% on Afri-JBB-Harm to 27.5%27.5\% on Afri-JBB-Culture and 25.7%25.7\% on AfriJail-Mono. Combined failure rate (ASR+Deflection\text{ASR} + \text{Deflection}) rises from 21.1%21.1\% on Afri-JBB-Harm to 43.2%43.2\% on Afri-JBB-Culture and 44.9%44.9\% on AfriJail-Mono.
    • For GPT-5.2, combined failure rate increases from 6.0%6.0\% on Afri-JBB-Harm to 17.2%17.2\% on Afri-JBB-Culture and 23.2%23.2\% on AfriJail-Mono.
    • For Grok-4.3, combined failure rate increases from 14.6%14.6\% on Afri-JBB-Harm to 21.0%21.0\% on Afri-JBB-Culture and 29.9%29.9\% on AfriJail-Mono.
    • For Claude-Haiku-4.5, combined failure rate rises from 43.1%43.1\% on Afri-JBB-Harm to 61.0%61.0\% on Afri-JBB-Culture and 55.3%55.3\% on AfriJail-Mono.
  6. Knowl 6 — Resource Hierarchy and Script Tokenization Impact on LLM Comprehension and Deflection

    empirical result

    LLM comprehension and safety behaviors across African languages vary systematically according to training resource tier and writing script:

    1. Latin-Script Resource Gradient: Among Latin-script African languages, safety refusal and comprehension track language resource tiers. Kiswahili (the highest-resourced) demonstrates the highest explicit refusal rates and lowest deflection rates across models. Hausa (second tier) displays comparable robustness. isiXhosa occupies an intermediate tier, while lower-resource languages—Igbo, Chichewa (Nyanja), and Yorùbá—suffer the highest deflection rates. For example, on AfriJail-Mono, Claude-Sonnet-5 Deflection is 24.5%24.5\% in Swahili, 18.1%18.1\% in Hausa, 32.9%32.9\% in Chichewa, 49.0%49.0\% in Igbo, and 29.7%29.7\% in Yorùbá.
    2. Script Tokenization Penalty in Amharic: Amharic diverges from this resource gradient. Despite having pretraining resources comparable to Hausa, Amharic exhibits consistently higher Deflection rates across evaluated models. Under GPT-4o on AfriJail-Mono, Deflection in Amharic is 43.7%43.7\% versus 14.9%14.9\% in Hausa; under Grok-3, it reaches 50.1%50.1\% in Amharic versus 19.2%19.2\% in Hausa; under Claude-Sonnet-5, it is 42.0%42.0\% in Amharic versus 18.1%18.1\% in Hausa. This elevated comprehension failure is attributed to tokenizer inefficiencies for the non-Latin Ge'ez script ("token tax"), which excessively fragments subword tokens and degrades semantic comprehension.
  7. Knowl 7 — Behavioral Effects of Boundary Point Jailbreaking and Code-Switching

    empirical result

    Adversarial evaluation via Boundary Point Jailbreaking (BPJ) and intra-sentential code-switching reveals distinct effects on safety guardrails versus comprehension bottlenecks:

    1. Boundary Point Jailbreaking (BPJ): BPJ introduces character perturbations across descending noise levels (20,15,12,9,6,3,020, 15, 12, 9, 6, 3, 0), stopping at the highest noise level that elicits a Jailbroken classification from the safety judge. BPJ significantly increases ASR relative to Direct Prompting across models (e.g., on Afri-JBB-Harm, average African-language ASR increases by +21.5+21.5 percentage points for GPT-4o, +18.9+18.9 for GPT-5.2, +21.0+21.0 for Grok-3, +18.3+18.3 for Grok-4.3, +13.5+13.5 for Claude-Haiku-4.5, and +23.0+23.0 for DeepSeek-3.2). BPJ partially reduces Deflection by converting off-target responses into jailbroken completions, but does not eliminate comprehension failures.
    2. Code-Switching (AfriJail-CS): Blending English with African languages reliably decreases Deflection across models relative to monolingual prompts (AfriJail-Mono), as English syntactic and semantic fragments assist prompt parsing. Average Deflection decreases by 8.08.0 percentage points in GPT-4o, 5.15.1 in GPT-5.2, 5.05.0 in Grok-4.3, 15.415.4 in Claude-Haiku-4.5, 10.210.2 in Llama-4-Maverick, and 9.19.1 in DeepSeek-3.2. However, ASR does not increase uniformly: improved comprehension manifests as either explicit refusal or jailbroken compliance depending on the model's underlying safety guardrails.
  8. Knowl 8 — Degradation of LLM-as-a-Judge Agreement in Low-Resource and Non-Latin Script African Languages

    empirical result

    Validation of LLM safety judges against human ground truth—constructed from majority voting across three independent native speakers per language evaluating 50 prompts on five target models (GPT-4o, GPT-5.2, Grok-3, Grok-4, DeepSeek-3.2) across six languages (1,5001,500 responses total)—reveals that LLM judge reliability deteriorates in low-resource and non-Latin script settings.

    Average agreement rates between LLM judges and human majority reference labels under Direct Prompting:

    LLM Judge Amharic (amh) Hausa (hau) Igbo (ibo) Chichewa (nya) Kiswahili (swh) Yorùbá (yor) Average
    GPT-4.1 57.0% 69.5% 57.8% 66.1% 79.8% 60.6% 65.1%
    Gemini 3.5 62.6% 75.7% 64.7% 71.2% 75.3% 67.3% 69.5%
    Claude Opus 5 60.1% 73.6% 67.1% 67.6% 72.0% 69.0% 68.2%

    Swahili achieves the highest judge-human agreement (75.3%–79.8%75.3\%\text{--}79.8\%), whereas lower-resource Latin languages (Igbo, Yorùbá) achieve only 57.8%–69.0%57.8\%\text{--}69.0\%. Amharic yields the lowest overall judge agreement (57.0%–62.6%57.0\%\text{--}62.6\%) and the highest incidence of three-way human annotator ties (3636 out of 250250 instances, or 14.4%14.4\%, compared to 44 in Igbo and 88 in Hausa), indicating that both target models and LLM evaluators struggle with comprehension and classification ambiguity under the Ge'ez script.

  9. Knowl 9 — Over-Refusal and Deflection Dynamics on Benign African Prompts

    empirical result

    Evaluating language models on 100 non-adversarial, benign prompts (Afri-JBB-Benign) under Direct Prompting reveals significant baseline over-refusal and elevated cross-lingual comprehension failure:

    • Over-Refusal: Models frequently refuse benign prompts across all languages, with refusal rates often exceeding 40%–60%40\%\text{--}60\% (e.g., GPT-5.2 over-refuses 72.0%72.0\% in English and 60.6%60.6\% on average across African languages; Grok-4.3 over-refuses 52.0%52.0\% in English and 44.7%44.7\% in African languages; Claude-Opus-4.8 over-refuses 46.0%46.0\% in English and 42.3%42.3\% in African languages). This suggests that safety guardrails heavily pattern-match on surface-level sensitive keywords rather than discerning benign intent.
    • Deflection on Benign Inputs: While compliance and refusal rates show no consistent separation between English and African languages, Deflection rates increase substantially in African languages. For GPT-4o, Deflection rises from 13.0%13.0\% in English to an African average of 25.1%25.1\%; for Grok-4.3, from 7.0%7.0\% to 14.6%14.6\%; for Claude-Haiku-4.5, from 12.0%12.0\% to 35.0%35.0\%; and for DeepSeek-3.2, from 11.0%11.0\% to 29.3%29.3\%. This confirms that low-resource language comprehension failure persists even when prompts lack any adversarial intent.
  10. Knowl 10 — Scope and Methodological Limitations of TukaBench

    limitation

    The authors identify four structural limitations of the TukaBench benchmark and its evaluation:

    1. Linguistic Scope: Although the benchmark covers seven African languages across the Niger-Congo and Afro-Asiatic families, it does not represent the full diversity of over 2,000 African languages or other major continental language families such as Nilo-Saharan or Khoisan.
    2. Exclusion of Extremely Low-Resource Languages: To ensure reliable native annotator recruitment, quality control, and human verification, the study restricts coverage to languages spoken by over 10 million people, potentially underestimating the severity of comprehension failures in even lower-resourced languages.
    3. Temporal Model Coverage: Evaluated LLMs represent a fixed snapshot of commercial and open-weight models, whose performance and alignment behaviors shift rapidly with new releases.
    4. Incomplete Model Results Due to Deprecation: Grok-3 and Grok-4 were deprecated by the provider during the experimental phase, resulting in missing evaluations for isiXhosa for those specific models.

Coverage note — None was omitted; all key contributions including dataset design, translation pipeline, evaluation framework, empirical results across attacks and models, human verification findings, and limitations are fully covered.

References

  1. 1.Tassallah Abdullahi, Macton Mgonzo, Mardiyyah Oduwole, Paul Okewunmi, Abraham Toluwase Owodunni, Ritambhara Singh, and Carsten Eickhoff. 2026. UbuntuGuard: A culturally-grounded policy benchmark for equitable AI safety in African languages. In Findings of the Association for Computational Linguistics: ACL 2026, pages 33262–33276, San Diego, California, United States. Association for Computational Linguistics.
  2. 2.David Ifeoluwa Adelani, Jesujoba Oluwadara Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Chinenye Emezue, Colin Leong, Michael Beukman, Shamsuddeen H. Muhammad, Guyo D. Jarso, Oreen Yousuf, and 26 others. 2022. A few thousand translations go a long way! Leveraging pre-trained models for African news translation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3053–3070, Seattle, United States. Association for Computational Linguistics.
  3. 3.Samee Arif, Naihao Deng, Zhijing Jin, and Rada Mihalcea. 2026. One word at a time: Incremental completion decomposition breaks llm safety. ArXiv, abs/2604.25921.
  4. 4.Berk Atil, R. Passonneau, and Fred Morstatter. 2025. Do methods to jailbreak and defend llms generalize across languages? ArXiv, abs/2511.00689.
  5. 5.Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, Hoda Heidari, Anson Ho, Sayash Kapoor, Leila Khalatbari, Shayne Longpre, Sam Manning, Vasilios Mavroudis, Mantas Mazeika, Julian Michael, and 77 others. 2025. International ai safety report. ArXiv, abs/2501.17805.
  6. 6.Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong. 2024. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. In Advances in Neural Information Processing Systems, volume 37, pages 55005–55029. Curran Associates, Inc.
  7. 7.Xander Davies, Giorgi Giglemiani, Edmund Lau, Eric Winsor, Geoffrey Irving, and Yarin Gal. 2026. Boundary point jailbreaking of black-box LLMs. arXiv preprint arXiv:2602.15001.
  8. 8.Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Lidong Bing. 2024. Multilingual jailbreak challenges in large language models. In International Conference on Learning Representations, volume 2024, pages 24634–24651.
  9. 9.Felix Friedrich, Simone Tedeschi, Patrick Schramowski, Manuel Brack, Roberto Navigli, Huu Nguyen, Bo Li, and Kristian Kersting. 2025. LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies. In ICLR 2025 Workshop on Building Trust in Language Models and Applications.
  10. 10.Rishav Hada, Varun Gumma, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2024a. METAL: Towards multilingual meta-evaluation. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2280–2298, Mexico City, Mexico. Association for Computational Linguistics.
  11. 11.Rishav Hada, Varun Gumma, Adrian de Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, and Sunayana Sitaram. 2024b. Are large language model-based evaluators the solution to scaling up multilingual evaluation? In Findings of the Association for Computational Linguistics: EACL 2024, pages 1051–1070, St. Julian’s, Malta. Association for Computational Linguistics.
  12. 12.Yutao Hou, Yihan Jiang, Yuhan Xie, Jian Yang, Liwen Zhang, Hailiang Huang, Guanhua Chen, and Yun Chen. 2026. FinSafetyBench: Evaluating LLM safety in real-world financial scenarios. In Findings of the Association for Computational Linguistics: ACL 2026, pages 14181–14208, San Diego, California, United States. Association for Computational Linguistics.
  13. 13.Devang Kulshreshtha, Hang Su, Haibo Jin, Chinmay Hegde, and Haohan Wang. 2026. Break me if you can: Self-jailbreaking of aligned LLMs via lexical insertion prompting. arXiv preprint arXiv:2601.02670.
  14. 14.Priyanshu Kumar, Devansh Jain, Akhila Yerukola, Liwei Jiang, Himanshu Beniwal, Thomas Hartvigsen, and Maarten Sap. 2025. Polyguard: A multilingual safety moderation tool for 17 languages. In Second Conference on Language Modeling.
  15. 15.Jie Li, Yi Liu, Chongyang Liu, Ling Shi, Xiaoning Ren, Yaowen Zheng, Yang Liu, and Yinxing Xue. 2024. A cross-language investigation into jailbreak attacks in large language models. ArXiv, abs/2401.16765.
  16. 16.Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, and David Ifeoluwa Adelani. 2025. SSACOMET: Do LLMs outperform learned metrics in evaluating MT for under-resourced African languages? In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 12979–12998, Suzhou, China. Association for Computational Linguistics.
  17. 17.Jessica M. Lundin, Ada Zhang, Nihal Karim, Hamza Louzan, Guohao Wei, David Ifeoluwa Adelani, and Cody Carroll. 2026. The token tax: Systematic bias in multilingual tokenization. In Proceedings of the 7th Workshop on African Natural Language Processing (AfricaNLP 2026), pages 103–112, Rabat, Morocco. Association for Computational Linguistics.
  18. 18.Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks. 2024. HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 35181–35224. PMLR.
  19. 19.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  20. 20.Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2685–2702, Online. Association for Computational Linguistics.
  21. 21.Jiayang Song, Yuheng Huang, Zhehua Zhou, and Lei Ma. 2025. Multilingual blending: Large language model safety alignment evaluation with language mixture. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 3433–3449, Albuquerque, New Mexico. Association for Computational Linguistics.
  22. 22.Xiaobing Sun, Perry Lam, Shaohua Li, Zizhou Wang, Rick Siow Mong Goh, Yong Liu, and Liangli Zhen. 2026. Structured semantic cloaking for jailbreak attacks on large language models. arXiv preprint arXiv:2603.16192.
  23. 23.Simone Tedeschi, Felix Friedrich, Patrick Schramowski, Kristian Kersting, Roberto Navigli, Huu Nguyen, and Bo Li. 2024. Alert: A comprehensive benchmark for assessing large language models’ safety through red teaming. ArXiv, abs/2404.08676.
  24. 24.Zheng Xin Yong, Cristina Menghini, and Stephen Bach. 2023. Low-resource languages jailbreak GPT-4. In Socially Responsible Language Modelling Research.
  25. 25.Haneul Yoo, Yongjin Yang, and Hwaran Lee. 2025. Code-switching red-teaming: LLM evaluation for safety and multilingual understanding. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13392–13413, Vienna, Austria. Association for Computational Linguistics.
  26. 26.Hao Yu, Tianyi Xu, Michael A. Hedderich, Wassim Hamidouche, Syed Waqas Zamir, and David Ifeoluwa Adelani. 2026. AfriqueLLM: How data mixing and model architecture impact continued pretraining for African languages. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5909–5928, San Diego, California, United States. Association for Computational Linguistics.
  27. 27.Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2024. SafetyBench: Evaluating the safety of large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15537–15553, Bangkok, Thailand. Association for Computational Linguistics.
  28. 28.Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Yu-Gang Jiang, and Bo Li. 2026. ML-Bench&Guard: Policy-grounded multilingual safety benchmark and guardrail for large language models. arXiv preprint arXiv:2605.00689.
  29. 29.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems, volume 36, pages 46595–46623. Curran Associates, Inc.
  30. 30.Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. Preprint, arXiv:2307.15043.

Citation

MLA
Akinode, V., et al. “TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages”. arXiv, 2026, http://arxiv.org/abs/2606.01322v2.
APA
Akinode, V., Li, S., Hamidouche, W., Zamir, W., Becker-Reshef, I., & Adelani, D. I. (2026). TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages. arXiv. http://arxiv.org/abs/2606.01322v2
Chicago
Akinode, V., S. Li, W. Hamidouche, W. Zamir, I. Becker-Reshef, and D. I. Adelani. 2026. “TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages”. arXiv. http://arxiv.org/abs/2606.01322v2.
Harvard
Akinode, V. et al. (2026) “TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2606.01322v2.
Vancouver
1. Akinode V, Li S, Hamidouche W, Zamir W, Becker-Reshef I, Adelani DI (2026) TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages. arXiv

BibTeX

@article{akinode2026tukabench,
  title = {TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages},
  author = {Akinode, Victor and Li, Senyu and Hamidouche, Wassim and Zamir, Waqas and Becker-Reshef, Inbal and Adelani, David Ifeoluwa},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2606.01322v2},
  eprint = {2606.01322}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/