Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate
Hannah KirkBertie VidgenPaul RöttgerTristan ThrushScott Hale
Introduces the HatemojiCheck diagnostic suite and the adversarially generated HatemojiBuild dataset to expose and correct critical blind spots in hate speech detection systems handling emoji-based abuse.
Online hate speech is a major societal challenge that harms individuals and degrades digital discourse, requiring automated moderation systems to handle large volumes of content. However, perpetrators increasingly use emoji to evade detection by substituting characters, replacing identity terms, or expressing threatening sentiments pictorially. Standard hate detection systems are rarely evaluated on or trained with these non-textual elements, leaving critical blind spots that allow toxic content to spread unaddressed.
The article evaluates how well current automated systems detect emoji-based hate speech and demonstrates how human-in-the-loop adversarial training can resolve model vulnerabilities.
To conduct this evaluation, the researchers created a functional test suite of 3,930 short statements across seven distinct emoji-use categories and six protected identities, pairing original hateful statements with minimally altered, non-hateful contrast cases. To address detected flaws, they then deployed an adversarial data generation framework across three iterative rounds. A team of trained human annotators generated 5,912 challenging, balanced examples designed to fool target machine learning models, retraining the models after each round to build stronger defenses.
The investigation revealed several key findings regarding model vulnerabilities and remediation. First, existing commercial and academic models fail substantially when faced with emoji-based hate. For instance, Google Jigsaw's Perspective API achieved only 68.9% accuracy on the functional test suite, and baseline models failed almost completely on cases where emoji replaced protected group terms. Second, incorporating dynamic adversarial training dramatically improved detection capabilities, raising overall test accuracy to nearly 88% and lifting performance on adversarial emoji test sets from an F1-score of 0.49 to over 0.76. Third, performance gains occurred rapidly, with a single round of roughly 2,000 adversarial examples driving the vast majority of improvement before returns plateaued. Finally, these improvements occurred without degrading performance on text-only hate speech or compromising fairness across demographic subgroups.
These results indicate that automated moderation systems currently deployed by major platforms possess severe, exploitable weaknesses against visual evasion tactics. Relying on standard text-focused training leaves organizations vulnerable to compliance failures, brand reputation damage, and user safety risks. Importantly, the findings demonstrate that fixing these vulnerabilities does not require massive compute or costly architectural overhauls; targeted, data-centric interventions can rapidly close performance gaps at modest expense.
Organizations developing or deploying content moderation tools should immediately integrate granular functional testing to audit their systems for emoji-based evasion before deployment. Engineering teams should adopt iterative, adversarial data collection pipelines rather than relying solely on static historical datasets. Furthermore, developers should prioritize robust tokenization strategies, such as Byte-Pair Encoding, rather than naive text translations that strip out subtle emoji context.
While the findings demonstrate high confidence in addressing targeted evasion strategies, decision-makers should note certain limitations. The evaluation suite relies on hand-crafted, short English-language statements covering six protected groups, meaning it establishes a minimum performance standard rather than full coverage of all real-world nuance. Future initiatives should expand these dynamic testing and training methodologies to multilingual contexts, intersectional identities, and emerging visual slang.
- Paper: Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, Marco Túlio Ribeiro et al. (2020). Introduces CheckList, the foundational behavioral testing and functional capability test suite framework upon which HatemojiCheck is directly conceptually structured and adapted.
- Paper: Automated Hate Speech Detection and the Problem of Offensive Language, Thomas Davidson et al. (2017). Establishes standard dataset construction and baseline distinctions between hate speech and general offensive language in social media moderation.
- Paper: Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter, Zeerak Waseem et al. (2016). Provides foundational methodology and feature analysis for automated social media hate speech detection.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Provides the background benchmark and methodology for evaluating toxicity generation and weaknesses in neural language systems.
- Paper: Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment, Di Jin et al. (2019). Demonstrates adversarial perturbation techniques that expose robustness failures in NLP classifiers, motivating human-and-model-in-the-loop adversarial generation.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). Extends safety classification and content moderation to conversational LLMs using adaptable, safety-aligned guardrail taxonomies.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). Builds upon adversarial dataset construction by automating the red teaming process with language models to expose toxic and harmful outputs at scale.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). Standardizes adversarial red teaming and refusal benchmarks across diverse models and multimodal safety settings.
- Paper: Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network, Bin Liang et al. (2022). Applies multimodal sentiment and figurative language detection to resolve complex contradictory visual-textual online posts.
