TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages
Victor AkinodeSenyu LiWassim HamidoucheWaqas ZamirInbal Becker-ReshefDavid Ifeoluwa Adelani
Introduces TukaBench to evaluate large language model safety across seven African languages, demonstrating that cultural context and code-switching significantly reduce model refusals while exposing critical failures in automated safety judging.
Safety evaluations for large language models remain heavily centered on English and a small group of high-resource languages. This leaves low-resource languages, especially African languages spoken by tens of millions of people, critically vulnerable to adversarial manipulation and jailbreaking attacks. Standard safety benchmarks typically rely on direct translations of Western-centric scenarios, failing to capture local cultural nuances and regional harm patterns.
The article evaluates model safety in low-resource environments by introducing TukaBench, the first culturally grounded jailbreak benchmark spanning English and seven African languages: Amharic, Hausa, Igbo, Chichewa, Kiswahili, Yorùbá, and isiXhosa. The benchmark tests how language translation, local cultural context, code-switching with English, and adversarial perturbation affect safety guardrails across thirteen leading proprietary and open-weight language models.
The benchmark encompasses 986 prompts per language across five datasets, constructed via machine translation followed by native-speaker post-editing and cultural adaptation. Prompts span ten misuse categories and incorporate culturally specific entities and harms, such as local financial fraud and governance issues. To evaluate responses, the authors introduce a three-part classification scheme: Refused, Jailbroken, and Deflected. The deflection metric identifies cases where a model neither complies with nor explicitly refuses a harmful prompt, but instead fails to comprehend the input and responds off-target. Automated evaluations using an automated model judge were validated against human majority labels across 1,500 model responses.
The evaluation reveals several critical findings regarding model safety in under-resourced languages. First, prompting in African languages reduces explicit model refusals from an average of 78.1% in English to 62.7%, while deflection jumps from 6.1% to 20.1%. Second, cultural grounding significantly increases model vulnerability: culturally adapted prompts generate higher failure rates than direct translations, proving that direct translation benchmarks underestimate real-world deployment risks. Third, language resource availability and script type strongly dictate comprehension; higher-resource languages like Swahili demonstrate stronger refusal behavior, whereas lower-resource languages and non-Latin scripts like Amharic suffer from severe comprehension breakdowns. Fourth, automated evaluation reliability drops sharply in lower-resource settings, where agreement between automated judges and native human annotators falls from roughly 80% in Swahili to below 60% in Yorùbá, Igbo, and Amharic. Finally, newer model generations and adversarial attacks like boundary point jailbreaking reduce off-target deflections, but often convert these newly understood prompts into harmful completions.
These findings indicate that relying strictly on attack success rates produces a false sense of security. Low attack rates in African languages often reflect poor language comprehension rather than robust safety alignment. Furthermore, automated safety evaluation pipelines become unreliable in low-resource and non-Latin script contexts, compounding deployment risks in multilingual regions. As models improve at language comprehension, their vulnerability to harmful compliance increases unless safety alignment is trained alongside linguistic capabilities.
Organizations developing or deploying models in multilingual environments must avoid using simple direct translation for safety auditing and instead adopt culturally localized test suites. Evaluation frameworks should incorporate deflection metrics to decouple comprehension errors from genuine refusals, and automated judging should be paired with human native-speaker spot checks. Because the benchmark focuses on relatively resourced African languages and snapshot model versions, future work must expand coverage to lower-resource dialects and build dedicated multilingual safety guardrails.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). This paper establishes the foundational concept of mismatched generalization, explaining how safety training fails when language models are prompted across languages and modalities where alignment was insufficiently applied.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This work introduces the LLM-as-a-judge paradigm and evaluates human agreement dynamics, providing essential context for understanding the automated evaluation failures in low-resource languages analyzed in TukaBench.
- Paper: Unintended Impacts of LLM Alignment on Global Representation, Michael J. Ryan et al. (2024). This study analyzes how standard alignment procedures introduce disparities across regional dialects and global cultural representations, providing direct background for cultural grounding in safety evaluations.
- Paper: MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks, Sanchit Ahuja et al. (2024). This cross-lingual benchmark highlights severe model performance degradation across low-resource African and Indic languages, contextualizing the language comprehension barriers encountered during safety benchmarking.
- Paper: The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants, Lucas Bandarkar et al. (2024). This paper establishes baseline multilingual comprehension capabilities across diverse scripts and low-resource language variants, clarifying the model comprehension limits observed in African language prompt evaluations.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). This work presents a standardized red-teaming framework and behavioral categorization for refusal and jailbreaks that informs the structured evaluation settings used in TukaBench.
- Paper: XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity, Dasol Choi et al. (2026). This benchmark broadens country-grounded adversarial safety and cultural sensitivity evaluation to a global multi-country setting across ten distinct national environments.
- Paper: HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment, Shei Pern Chua et al. (2026). This work introduces a geometric alignment defense that directly addresses the prompt-level refusal failures and evasion mechanisms exposed in multilingual and culturally grounded jailbreak benchmarks.
