Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate

Hannah KirkBertie VidgenPaul RöttgerTristan ThrushScott Hale

article2022NAACL85 citations

Introduces the HatemojiCheck diagnostic suite and the adversarially generated HatemojiBuild dataset to expose and correct critical blind spots in hate speech detection systems handling emoji-based abuse.

Listen

Online hate speech is a major societal challenge that harms individuals and degrades digital discourse, requiring automated moderation systems to handle large volumes of content. However, perpetrators increasingly use emoji to evade detection by substituting characters, replacing identity terms, or expressing threatening sentiments pictorially. Standard hate detection systems are rarely evaluated on or trained with these non-textual elements, leaving critical blind spots that allow toxic content to spread unaddressed.

The article evaluates how well current automated systems detect emoji-based hate speech and demonstrates how human-in-the-loop adversarial training can resolve model vulnerabilities.

To conduct this evaluation, the researchers created a functional test suite of 3,930 short statements across seven distinct emoji-use categories and six protected identities, pairing original hateful statements with minimally altered, non-hateful contrast cases. To address detected flaws, they then deployed an adversarial data generation framework across three iterative rounds. A team of trained human annotators generated 5,912 challenging, balanced examples designed to fool target machine learning models, retraining the models after each round to build stronger defenses.

The investigation revealed several key findings regarding model vulnerabilities and remediation. First, existing commercial and academic models fail substantially when faced with emoji-based hate. For instance, Google Jigsaw's Perspective API achieved only 68.9% accuracy on the functional test suite, and baseline models failed almost completely on cases where emoji replaced protected group terms. Second, incorporating dynamic adversarial training dramatically improved detection capabilities, raising overall test accuracy to nearly 88% and lifting performance on adversarial emoji test sets from an F1-score of 0.49 to over 0.76. Third, performance gains occurred rapidly, with a single round of roughly 2,000 adversarial examples driving the vast majority of improvement before returns plateaued. Finally, these improvements occurred without degrading performance on text-only hate speech or compromising fairness across demographic subgroups.

These results indicate that automated moderation systems currently deployed by major platforms possess severe, exploitable weaknesses against visual evasion tactics. Relying on standard text-focused training leaves organizations vulnerable to compliance failures, brand reputation damage, and user safety risks. Importantly, the findings demonstrate that fixing these vulnerabilities does not require massive compute or costly architectural overhauls; targeted, data-centric interventions can rapidly close performance gaps at modest expense.

Organizations developing or deploying content moderation tools should immediately integrate granular functional testing to audit their systems for emoji-based evasion before deployment. Engineering teams should adopt iterative, adversarial data collection pipelines rather than relying solely on static historical datasets. Furthermore, developers should prioritize robust tokenization strategies, such as Byte-Pair Encoding, rather than naive text translations that strip out subtle emoji context.

While the findings demonstrate high confidence in addressing targeted evasion strategies, decision-makers should note certain limitations. The evaluation suite relies on hand-crafted, short English-language statements covering six protected groups, meaning it establishes a minimum performance standard rather than full coverage of all real-world nuance. Future initiatives should expand these dynamic testing and training methodologies to multilingual contexts, intersectional identities, and emerging visual slang.

Cover for Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate

Abstract

Detecting online hate is a complex task, and low-performing models have harmful consequences when used for sensitive applications such as content moderation. Emoji-based hate is an emerging challenge for automated detection. We present HATEMOJICHECK, a test suite of 3,930 short-form statements that allows us to evaluate performance on hateful language expressed with emoji. Using the test suite, we expose weaknesses in existing hate detection models. To address these weaknesses, we create the HATEMOJIBUILD dataset using a human-and-model-in-the-loop approach. Models built with these 5,912 adversarial examples perform substantially better at detecting emoji-based hate, while retaining strong performance on text-only hate. Both HATEMOJICHECK and HATEMOJI-BUILD are made publicly available.

Table of Contents

  • 1 Introduction
  • 2 HATEMOJICHECK: Functional Tests for Emoji-Based Hate
  • 2.1 Identifying Functionalities
  • 2.2 Functionalities in HATEMOJICHECK
  • 2.3 Test Cases in HATEMOJICHECK
  • 2.3.1 Perturbations
  • 2.4 Validating Test Cases
  • 3 Building Better Models with HATEMOJIBUILD
  • 4 Evaluating Models with HATEMOJICHECK
  • 4.1 Model Performance by Functionality
  • 4.2 Model Performance by Target Group
  • 5 Discussion
  • 6 Related Work
  • 7 Conclusion
  • Acknowledgments
  • References
  • A Ethical Considerations
  • B Data Statement for HATEMOJICHECK
  • B.1 Curation Rationale
  • B.2 Language Variety
  • B.3 Speaker Demographics
  • B.4 Annotator Demographics
  • B.5 Speech Situation
  • B.6 Text Characteristics
  • C Constructing HATEMOJICHECK
  • C.1 List of Hateful Slurs in HATEMOJICHECK
  • C.2 Defining Perturbations
  • D Data Statement for HATEMOJIBUILD
  • D.1 Curation Rationale
  • D.2 Language Variety
  • D.3 Speaker Demographics
  • D.4 Annotator Demographics
  • D.5 Speech Situation
  • D.6 Text Characteristics
  • E Constructing HATEMOJIBUILD
  • F Target Models
  • G Robustness Analysis of Baselines
  • H Fairness Considerations

Knowls

  1. Knowl 1 — HATEMOJICHECK Functional Testing Suite for Emoji-Based Hate Speech

    model/method

    HATEMOJICHECK is a functional testing diagnostic suite consisting of 3,930 short-form English test cases designed to evaluate whether content moderation models correctly identify hate speech expressed with emoji. It comprises 2,126 original hateful test cases and 1,804 non-hateful perturbations across 88 unique templates, utilizing 36 emoji and populating placeholders with 6 protected identity groups (women, trans people, gay people, Black people, disabled people, and Muslims) as well as 16 target-specific slurs.

    The suite evaluates seven distinct functional categories of emoji usage:

    1. F1 (Verb Swap): Threatening hate speech where the violent verb is replaced with an emoji (e.g., violent actions represented by knife, gun, bomb, blood drop, or dagger emoji). Includes direct threats and normative threats (300 cases).
    2. F2 (Identity Swap): Derogatory hate speech where the protected identity term is substituted with an emoji representation (e.g., woman emoji for women, transgender flag for trans people, rainbow flag for gay people, wheelchair symbol for disabled people, monkey emoji for Black people, person wearing turban for Muslims) (120 cases).
    3. F3 (Descriptor Swap): Hate speech where negative descriptors, dehumanizing terms, or disgust expressions are replaced with emoji (e.g., rat, pig, pile of poo, cockroach, worm, banana, airplane, vomiting face, nauseated face, face with medical mask) (260 cases).
    4. F4 (Double Swap): Combines F1 with F2 or F3 by substituting emoji for both the identity and the verb/descriptor (288 cases).
    5. F5 (Append): Appends negative/hostile emoji (e.g., vomiting face, pile of poo, clown face, wastebasket, thumbs down) to an otherwise neutral statement, rendering the whole statement hateful (288 cases).
    6. F6 (Positive Confounder): Appends positive emoji (e.g., red heart, sparkles, smiling face with heart-eyes, party popper, rainbow) to explicit hateful text to evaluate if positive sentiment cues fool the classifier (440 cases).
    7. F7 (Emoji Leetspeak): Obfuscates hateful text, identities, or slurs by substituting characters or word-pieces with emoji, including numbers (e.g., digit four for 'a', digit one for 'i', digit three for 'e', digit zero for 'o'), star or heart vowel substitutions, and pictographic homophones/word pieces (430 cases).
  2. Knowl 2 — Perturbation Mechanisms in HATEMOJICHECK

    model/method

    To evaluate model decision boundaries with minimal lexical and syntactic variation, HATEMOJICHECK pairs original hateful test templates with three contrasting perturbation types, generating 1,804 perturbation test cases:

    1. Identity Perturbations (314 cases): The protected identity term or emoji in the original template is replaced by a non-protected entity, such as non-protected professions (e.g., "accountants"), non-human animals (e.g., "spiders"), or inanimate objects (e.g., "pizza"). This preserves the grammatical structure and sentiment but flips the ground-truth label to non-hateful.
    2. Polarity Perturbations (902 cases): The negative sentiment, threat, or derogatory term/emoji in the template is converted into positive or supportive language (e.g., replacing "kill" with "respect", or "I hate..." with "I love..."), or by transforming a slur usage into counterspeech (e.g., replacing "[IDENTITY] are [SLUR]" with "[IDENTITY] should never be called [SLUR]"). This preserves the target identity while flipping the ground-truth label to non-hateful.
    3. No-Emoji Perturbations (588 cases): All emoji tokens in the template are replaced with equivalent textual words (e.g., replacing gun emoji with "shoot", heart emoji with "love") or removed entirely (for appended emoji in F5 and F6). For functionalities F1–F4, F6, and F7, this keeps the label hateful; for F5 (Append), removing the appended negative emoji renders the remaining text non-hateful.
  3. Knowl 3 — HATEMOJIBUILD Adversarial Dataset Generation

    data/table

    HATEMOJIBUILD is a dynamic, human-and-model-in-the-loop dataset generated to train models on emoji-based hate speech. Spanning three consecutive rounds of collection (denoted R5, R6, and R7, continuing from four prior text-only rounds R1–R4), 10 trained human annotators were tasked with generating realistic, linguistically diverse entries containing at least one emoji that fool the target model in the loop.

    Each input was assigned a binary label (hate vs. not hate) and, if hateful, secondary labels for hate type (Derogation, Animosity, Threatening language, Dehumanizing language) and target identity. For every validated original entry, annotators constructed an offline perturbation flipping the binary label by modifying either the emoji while fixing the text or vice versa.

    Dataset Attribute Round 5 (R5) Round 6 (R6) Round 7 (R7)
    Total entries (nn) 1994 1966 1952
    Train split, nn (%) 1595 (80.0) 1572 (80.0) 1561 (80.0)
    Dev split, nn (%) 199 (10.0) 197 (10.0) 195 (10.0)
    Test split, nn (%) 200 (10.0) 197 (10.0) 196 (10.0)
    Hate, nn (%) 1006 (50.5) 983 (50.0) 976 (50.0)
    Not hate, nn (%) 988 (49.5) 983 (50.0) 976 (50.0)
    Original entries, nn (%) 997 (50.0) 983 (50.0) 976 (50.0)
    Perturbation entries, nn (%) 997 (50.0) 983 (50.0) 976 (50.0)
    Non-hate (None), nn (%) 988 (49.5) 983 (50.0) 976 (50.0)
    Derogation, nn (%) 718 (36.0) 649 (33.0) 594 (30.4)
    Animosity, nn (%) 74 (3.7) 219 (11.1) 275 (14.1)
    Threatening, nn (%) 101 (5.1) 50 (2.5) 52 (2.7)
    Dehumanizing, nn (%) 113 (5.7) 65 (3.3) 55 (2.8)
    Mean emoji count (μ±σ\mu \pm \sigma) 1.7±2.21.7 \pm 2.2 1.7±1.01.7 \pm 1.0 1.6±1.11.6 \pm 1.1

    In total, HATEMOJIBUILD contains 5,912 examples (50% hate, 50% non-hate) covering 54 unique targets and 126 intersectional identities, using 1,082 unique emoji tokens.

  4. Knowl 4 — Dynamic Model Retraining and Architecture Selection Scheme

    algorithm

    In each round of dynamic data generation for HATEMOJIBUILD, a new target model was selected from candidate architectures and upsampling configurations trained on all accumulated historical training data.

    Input: Historical training sets D0,…,Dt−1D_0, \dots, D_{t-1}, current round training set DtD_t, historical test sets Tprior=⋃i=1t−1TiT_{\text{prior}} = \bigcup_{i=1}^{t-1} T_i, current round test set TtT_t
    Output: Selected target model Mt∗M_t^* for round t+1t+1
    for each candidate architecture A∈{DeBERTa-base,BERTweet}A \in \{\text{DeBERTa-base}, \text{BERTweet}\} do
        for each upsampling multiplier u∈{1,5,10,100}u \in \{1, 5, 10, 100\} on DtD_t do
            Construct upsampled dataset D~t←u×Dt\tilde{D}_t \leftarrow u \times D_t
            Construct combined training set Dtrain←⋃i=0t−1D~i∪D~tD_{\text{train}} \leftarrow \bigcup_{i=0}^{t-1} \tilde{D}_i \cup \tilde{D}_t
            Train candidate model MA,uM_{A, u} on DtrainD_{\text{train}} for 3 epochs with early stopping on dev loss (learning rate 2×10−52 \times 10^{-5}, weighted Adam optimizer)
            Calculate prior test set accuracy: Accprior←Accuracy(MA,u,Tprior)\text{Acc}_{\text{prior}} \leftarrow \text{Accuracy}(M_{A, u}, T_{\text{prior}})
            Calculate current test set accuracy: Acccurrent←Accuracy(MA,u,Tt)\text{Acc}_{\text{current}} \leftarrow \text{Accuracy}(M_{A, u}, T_t)
            Compute selection metric: Score(MA,u)←0.5×Accprior+0.5×Acccurrent\text{Score}(M_{A, u}) \leftarrow 0.5 \times \text{Acc}_{\text{prior}} + 0.5 \times \text{Acc}_{\text{current}}
        end for
    end for
    Select Mt∗←arg⁡max⁡A,uScore(MA,u)M_t^* \leftarrow \arg\max_{A, u} \text{Score}(M_{A, u})
    return Mt∗M_t^*

    Across all rounds, DeBERTa (134M parameters, BPE tokenizer preserving emoji tokens) consistently outperformed BERTweet (135M parameters, text emoji translator via the Python emoji package). The selected target models were:

    • R6-T: DeBERTa trained with 100×100\times upsampling on R5.
    • R7-T: DeBERTa trained with 1×1\times upsampling on R6.
    • R8-T: DeBERTa trained with 5×5\times upsampling on R7.
  5. Knowl 5 — Model Error Rate Trajectory Across Dynamic Adversarial Rounds

    empirical result

    Model Error Rate (MER), defined as the percentage of original entries generated by annotators that successfully fool the model in the loop, reveals model susceptibility to emoji attacks:

    • Across text-only rounds (R1 to R4), MER steadily declined from approximately 55%55\% in R1 to under 30%30\% by R4 as the model adapted to textual adversarial patterns.
    • In R5, when annotators were first instructed to introduce emoji-based attacks against the text-trained DeBERTa target model (R5-T), the MER surged by 35 percentage points to over 60%60\%, surpassing the error rate of R1.
    • In R6, after retraining the target model on R5 emoji data (yielding model R6-T), MER dropped sharply by 24 percentage points to approximately 37%37\%.
    • In R7, annotators faced model R7-T and adjusted their strategies from direct emoji substitutions to irony, satire, and mockery, resulting in a slight further MER reduction of 1 percentage point to 36%36\%.

    Overall, the target model was easily fooled by emoji constructions upon initial exposure but became substantially more robust after a single round of adversarial retraining.

  6. Knowl 6 — Benchmark Performance of Pre-Emoji vs. Adversarially Trained Target Models

    data/table

    Evaluating pre-emoji baseline models versus adversarially retrained emoji-aware target models across emoji test sets (HATEMOJIBUILD R5–R7, HATEMOJICHECK), text test sets (R1–R4, HATECHECK), and combined datasets demonstrates that dynamic emoji training provides substantial gains on emoji hate without degrading performance on text-only hate.

    Emoji Test Sets Text Test Sets All Rounds
    R5–R7 (n=593n=593) HATEMOJICHECK (n=3930n=3930) R1–R4 (n=4119n=4119) HATECHECK (n=3728n=3728) R1–R7 (n=4712n=4712)
    Model Acc F1 Acc F1 Acc F1 Acc F1 Acc F1
    P-IA 0.508 0.394 0.689 0.754 0.679 0.720 0.765 0.839 0.658 0.689
    P-TX 0.523 0.448 0.650 0.711 0.602 0.659 0.720 0.813 0.592 0.639
    B-D 0.489 0.270 0.578 0.636 0.589 0.607 0.632 0.738 0.591 0.586
    B-F 0.496 0.322 0.552 0.605 0.562 0.562 0.602 0.694 0.557 0.532
    R5-T 0.585 0.490 0.779 0.825 0.828 0.847 0.956 0.968 0.786 0.801
    R6-T 0.757 0.769 0.879 0.910 0.823 0.837 0.961 0.971 0.813 0.825
    R7-T 0.759 0.762 0.867 0.899 0.824 0.842 0.955 0.967 0.813 0.829
    R8-T 0.744 0.755 0.871 0.904 0.827 0.844 0.966 0.975 0.814 0.829

    Pre-emoji models—including Google Perspective API Identity Attack (P-IA) and Toxicity (P-TX), academic BERT models trained on the Davidson et al. (B-D) and Founta et al. (B-F) datasets, and the text-trained DeBERTa baseline (R5-T)—perform poorly on emoji benchmarks (F1 scores between 0.2700.270 and 0.4900.490 on R5–R7). Training on just one round of adversarial emoji data (R6-T) yields an F1 gain of +0.279+0.279 on R5–R7 (0.490→0.7690.490 \to 0.769) and +0.100+0.100 accuracy on HATEMOJICHECK (0.779→0.8790.779 \to 0.879), with performance saturating in subsequent rounds (R7-T and R8-T).

  7. Knowl 7 — Model Diagnostic Accuracy Across HATEMOJICHECK Functionalities

    data/table

    Evaluating models across HATEMOJICHECK's seven functionalities and their associated perturbation subsets reveals specific vulnerabilities in baseline models that are resolved by dynamic emoji training.

    Functionality / Perturbation Label nn P-IA R5-T R6-T R7-T R8-T
    F1 Verb Swap hate 300 0.94 0.85 0.84 0.79 0.84
    F1.1 Identity perturbation not hate 50 0.94 1.00 1.00 0.98 0.86
    F1.2 Polarity perturbation not hate 60 0.20 0.42 0.75 0.77 0.73
    F1.3 No emoji perturbation hate 60 1.00 0.95 0.97 0.97 0.98
    F2 Identity Swap hate 120 0.14 0.33 0.83 0.83 0.83
    F2.1 Identity perturbation not hate 20 0.90 0.70 1.00 0.85 1.00
    F2.2 Polarity perturbation not hate 120 1.00 0.98 0.78 0.88 0.93
    F2.3 No emoji perturbation hate 120 0.98 0.98 0.97 1.00 0.95
    F3 Descriptor Swap hate 260 0.92 0.83 0.99 0.99 1.00
    F3.1 Identity perturbation not hate 40 1.00 1.00 1.00 1.00 1.00
    F3.2 Polarity perturbation not hate 60 0.25 0.48 0.78 0.82 0.93
    F3.3 No emoji perturbation hate 60 0.98 1.00 1.00 1.00 1.00
    F4 Double Swap hate 288 0.00 0.03 0.79 0.70 0.77
    F4.1 Identity perturbation not hate 46 1.00 1.00 0.91 0.98 0.91
    F4.2 Polarity perturbation not hate 60 1.00 0.98 0.92 0.85 0.92
    F4.3 No emoji perturbation hate 60 0.97 1.00 1.00 1.00 1.00
    F5 Append hate 288 0.73 0.69 0.99 0.89 0.87
    F5.1 Identity perturbation not hate 48 1.00 1.00 1.00 1.00 1.00
    F5.2 Polarity perturbation not hate 60 0.47 0.55 0.85 0.78 0.75
    F5.3 No emoji perturbation not hate 60 0.35 0.48 0.45 0.43 0.37
    F6 Positive Confounder hate 440 0.93 1.00 0.85 0.90 0.89
    F6.1 Identity perturbation not hate 65 0.75 0.92 0.92 0.92 0.92
    F6.2 Polarity perturbation not hate 112 0.44 0.88 0.96 0.91 0.93
    F6.3 No emoji perturbation hate 88 0.93 0.99 0.95 1.00 0.93
    F7 Emoji Leetspeak hate 430 0.59 0.85 0.91 0.83 0.91
    F7.1 Identity perturbation not hate 45 0.82 0.67 0.62 0.64 0.49
    F7.2 Polarity perturbation not hate 430 0.57 0.79 0.79 0.86 0.77
    F7.3 No emoji perturbation hate 140 0.61 0.99 0.99 0.94 0.99

    Key diagnostic findings include:

    • Identity Swaps (F2, F4): Pre-emoji models fail catastrophically when identity words are replaced with emoji (P-IA scores 0.140.14 on F2 and 0.000.00 on F4; R5-T scores 0.330.33 on F2 and 0.030.03 on F4). Emoji-aware target models achieve 0.830.83 on F2 and 0.700.70–0.790.79 on F4.
    • Polarity Inversion Vulnerability (F1.2, F3.2): Baselines overfit to identity words and predict hate despite positive wording (P-IA achieves 0.200.20 on F1.2 and 0.250.25 on F3.2; R5-T achieves 0.420.42 on F1.2 and 0.480.48 on F3.2). Retrained models improve accuracy up to 0.770.77 on F1.2 and 0.930.93 on F3.2.
    • Positive Confounders (F6): Baselines ignore appended positive emoji, maintaining high accuracy on hateful originals (0.930.93 for P-IA, 1.001.00 for R5-T) but failing on polarity-flipped statements (0.440.44 for P-IA on F6.2). Retrained target models achieve consistent accuracy (0.850.85–0.960.96) across both original and perturbed cases.
  8. Knowl 8 — Subgroup Fairness Disparities Across Target Identities

    empirical result

    Evaluating models across six protected identity subgroups (Muslims, Black people, disabled people, gay people, trans people, and women) in HATEMOJICHECK demonstrates severe subgroup disparities in commercial and static baselines.

    Model Demographic Parity Ratio Equalized Odds Ratio
    Perspective API Identity Attack (P-IA) 0.668 0.430
    Perspective API Toxicity (P-TX) 0.611 0.405
    BERT trained on Davidson et al. (B-D) 0.174 0.138
    BERT trained on Founta et al. (B-F) 0.276 0.257
    DeBERTa R5-T (Text-only baseline) 0.898 0.767
    DeBERTa R6-T (Target model Round 6) 0.818 0.711
    DeBERTa R7-T (Target model Round 7) 0.848 0.667
    DeBERTa R8-T (Target model Round 8) 0.832 0.600
    • Perspective API (P-IA): Exhibits unbalanced error distributions across groups, with high false positive rates for statements about women and high false negative rates for statements about disabled people.
    • Academic Baselines (B-D, B-F): Exhibit severe demographic disparity and equalized odds imbalance, with Equalized Odds Ratios below 0.260.26.
    • Adversarially Retrained Models: Although text-only model R5-T achieves a higher Demographic Parity Ratio (0.8980.898) and Equalized Odds Ratio (0.7670.767) due to uniform insensitivity to emoji across all groups, the emoji-aware model R8-T achieves strictly superior absolute performance—higher accuracy, precision, and recall, as well as lower false positive and false negative rates—across every individual identity subgroup.
  9. Knowl 9 — Limitations of HATEMOJICHECK and Dynamic Adversarial Emoji Datasets

    limitation

    The construction and evaluation of HATEMOJICHECK and HATEMOJIBUILD have four primary limitations:

    1. Negative Predictive Power of Diagnostic Functional Tests: HATEMOJICHECK relies on short, template-based, unambiguous statements. Achieving high performance on this test suite indicates solely the absence of specific minimal-case weaknesses, not robustness to naturalistic, complex, or long-form hate speech.
    2. Constrained Linguistic, Demographic, and Unicode Scope: The test suite is restricted to English, evaluates only 6 non-intersectional protected identities, and covers 36 emoji out of more than 3,500 defined in the Unicode Standard. Because emoji semantics and pragmatics vary across cultural and linguistic contexts, findings may not directly generalize to other languages or cultures.
    3. Synthetic Nature of Dynamic Adversarial Datasets: Data generated via human-and-model-in-the-loop procedures is synthetic rather than sampled from live social media platforms. Annotators can exhaust their creativity over successive rounds, producing simplistic or unnatural constructions, and the small annotator pool (11 annotators) may introduce idiosyncratic biases.
    4. Rapid Performance Saturation: Model performance gains on functional diagnostic tests saturate quickly after a single round of adversarial training (R5→R6R5 \to R6), offering limited incremental improvements in subsequent collection rounds.

Coverage note — None was omitted; all primary contributions, benchmark definitions, dataset statistics, training algorithms, empirical results across test sets/functionalities, fairness evaluations, and limitations are fully covered.

References

  1. 1.Alekh Agarwal, Alina Beygelzimer, Miroslav Dud<unk>ık, John Langford, and Hanna Wallach. 2018. A Reductions Approach to Fair Classification. arXiv:1803.02453 [cs].
  2. 2.Alekh Agarwal, Miroslav Dudik, and Zhiwei Steven Wu. 2019. Fair Regression: Quantitative Definitions and Reduction-Based Algorithms. In Proceedings of the 36th International Conference on Machine Learning, pages 120–129. PMLR.
  3. 3.Francesco Barbieri, German Kruszewski, Francesco Ronzano, and Horacio Saggion. 2016. How Cosmopolitan Are Emojis? Exploring Emojis Usage and Meaning over Different Languages with Distributional Semantics. In Proceedings of the 24th ACM International Conference on Multimedia, MM ’16, pages 531–535, New York, NY, USA. Association for Computing Machinery.
  4. 4.Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp. 2020. Beat the AI: Investigating adversarial human annotation for reading comprehension. Transactions of the Association for Computational Linguistics, 8:662–678.
  5. 5.Anna Bax. 2018. “The C-Word” Meets “the N-Word”: The Slur-Once-Removed and the Discursive Construction of “Reverse Racism”. Journal of Linguistic Anthropology, 28(2):114–136.
  6. 6.Boris Beizer. 1995. Black-Box Testing: Techniques for Functional Testing of Software and Systems. John Wiley & Sons, Inc., New York, NY, USA.
  7. 7.Emily M. Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  8. 8.Eckhard Bick. 2020. Annotating emoticons and emojis in a german-danish social media corpus for hate speech research. RASK, 52.
  9. 9.Brandwatch. 2018. The Emoji Report.
  10. 10.Spencer Cappallo, Stacey Svetlichnaya, Pierre Garrigues, Thomas Mensink, and Cees G. M. Snoek. 2019. New Modality: Emoji Challenges in Prediction, Anticipation, and Retrieval. IEEE Transactions on Multimedia, 21(2):402–415.
  11. 11.Michele Corazza, Stefano Menini, Elena Cabrio, Sara Tonelli, and Serena Villata. 2020. Hybrid Emoji-Based Masked Language Models for Zero-Shot Abusive Language Detection. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 943–949, Online. Association for Computational Linguistics.
  12. 12.Juliet M. Corbin and Anselm Strauss. 1990. Grounded theory research: Procedures, canons, and evaluative criteria. Qualitative Sociology, 13(1):3–21.
  13. 13.Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated Hate Speech Detection and the Problem of Offensive Language. Proceedings of the International AAAI Conference on Web and Social Media, 11(1):512–515.
  14. 14.Leon Derczynski, Hannah Rose Kirk, Abeba Birhane, and Bertie Vidgen. 2022. Handling and Presenting Harmful Text. arXiv:2204.14256 [cs].
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  16. 16.Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019. Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4537–4546, Hong Kong, China. Association for Computational Linguistics.
  17. 17.Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018. Measuring and Mitigating Unintended Bias in Text Classification. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, AIES ’18, pages 67–73, New York, NY, USA. Association for Computing Machinery.
  18. 18.Bjarke Felbo, Alan Mislove, Anders Søgaard, Iyad Rahwan, and Sune Lehmann. 2017. Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1615–1625, Copenhagen, Denmark. Association for Computational Linguistics.
  19. 19.Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis. 2018. Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior. In Twelfth International AAAI Conference on Web and Social Media.
  20. 20.Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou. 2020. Evaluating Models’ Local Decision Boundaries via Contrast Sets. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1307–1323, Online. Association for Computational Linguistics.
  21. 21.Katharine Gelber and Luke McNamara. 2016. Evidencing the harms of hate speech. Social Identities, 22(3):324–341.
  22. 22.Tarleton Gillespie. 2020. Content moderation, AI, and the question of scale. Big Data and Society, 7(2).
  23. 23.Moritz Hardt, Eric Price, and Nati Srebro. 2016. Equality of Opportunity in Supervised Learning. In Advances in Neural Information Processing Systems, volume 29. Curran Associates, Inc.
  24. 24.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. DeBERTa: Decoding-enhanced BERT with Disentangled Attention. arXiv:2006.03654 [cs].
  25. 25.Muhammad Okky Ibrohim, Muhammad Akbar Setiadi, and Indra Budi. 2019. Identification of hate speech and abusive language on indonesian Twitter using the Word2vec, part of speech and emoji features. In Proceedings of the International Conference on Advanced Information Science and System, AISS ’19, pages 1–5, New York, NY, USA. Association for Computing Machinery.
  26. 26.Alastair Jamieson. 2020. Racist comments after Euro 2020: Saka, Sancho and Rashford racially abused online after England defeat.
  27. 27.Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021. Dynabench: Rethinking Benchmarking in NLP. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4110–4124, Online. Association for Computational Linguistics.
  28. 28.J. Richard Landis and Gary G. Koch. 1977. The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1):159–174.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs].
  30. 30.Nikola Ljubešić and Darja Fišer. 2016. A global analysis of emoji usage. In Proceedings of the 10th Web as Corpus Workshop, pages 82–89, Stroudsburg, PA, USA. Association for Computational Linguistics.
  31. 31.Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, and Douwe Kiela. 2021. Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking. In Advances in Neural Information Processing Systems, volume 34, pages 10351–10367. Curran Associates, Inc.
  32. 32.Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. 2020. BERTweet: A pre-trained language model for English Tweets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 9–14, Online. Association for Computational Linguistics.
  33. 33.Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2021. DynaSent: A Dynamic Benchmark for Sentiment Analysis. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2388–2404, Online. Association for Computational Linguistics.
  34. 34.Justas Randolph. 2005. Free-marginal multirater kappa: An alternative to fleiss. In Joensuu Learning and Instruction Symposium, page 20.
  35. 35.Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond Accuracy: Behavioral Testing of NLP Models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902–4912, Online. Association for Computational Linguistics.
  36. 36.David Rodrigues, Marília Prada, Rui Gaspar, Margarida V. Garrido, and Diniz Lopes. 2018. Lisbon Emoji and Emoticon Database (LEED): Norms for emoji and emoticons in seven evaluative dimensions. Behavior Research Methods, 50(1):392–405.
  37. 37.Paul Röttger, Bertie Vidgen, Dirk Hovy, and Janet B. Pierrehumbert. 2022. Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks. arXiv:2112.07475 [cs].
  38. 38.Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021. HateCheck: Functional Tests for Hate Speech Detection Models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 41–58, Online. Association for Computational Linguistics.
  39. 39.Niloofar Safi Samghabadi, Afsheen Hatami, Mahsa Shafaei, Sudipta Kar, and Thamar Solorio. 2019. Attending the emotions to detect online abusive language.
  40. 40.Leandro Silva, Mainack Mondal, Denzil Correa, Fabrício Benevenuto, and Ingmar Weber. 2016. Analyzing the targets of hate in online social media. arXiv.
  41. 41.Zeerak Talat, Thomas Davidson, Dana Warmsley, and Ingmar Weber. 2017. Understanding Abuse: A Typology of Abusive Language Detection Subtasks. In Proceedings of the First Workshop on Abusive Language Online, pages 78–84, Vancouver, BC, Canada. Association for Computational Linguistics.
  42. 42.United Nations. 2019. UN strategy and plan of action on hate speech. Technical report.
  43. 43.Bertie Vidgen, Alex Harris, Dong Nguyen, Rebekah Tromble, Scott Hale, and Helen Margetts. 2019. Challenges and frontiers in abusive content detection. In Proceedings of the Third Workshop on Abusive Language Online, pages 80–93, Florence, Italy. Association for Computational Linguistics.
  44. 44.Bertie Vidgen, Dong Nguyen, Helen Margetts, Patricia Rossini, and Rebekah Tromble. 2021a. Introducing CAD: The Contextual Abuse Dataset. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2289–2303, Online. Association for Computational Linguistics.
  45. 45.Bertie Vidgen, Tristan Thrush, Zeerak Waseem, and Douwe Kiela. 2021b. Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1667–1682, Online. Association for Computational Linguistics.
  46. 46.Michael Wiegand and Josef Ruppenhofer. 2021. Exploiting Emojis for Abusive Language Detection. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 369–380, Online. Association for Computational Linguistics.
  47. 47.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. HuggingFace’s Transformers: State-of-the-art Natural Language Processing. arXiv:1910.03771 [cs].
  48. 48.Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. 2019. SemEval-2019 Task 6: Identifying and Categorizing Offensive Language in Social Media (OffensEval). In Proceedings of the 13th International Workshop on Semantic Evaluation, pages 75–86, Minneapolis, Minnesota, USA. Association for Computational Linguistics.
  49. 49.Xiang Zhou, Yixin Nie, Hao Tan, and Mohit Bansal. 2020. The Curse of Performance Instability in Analysis Datasets: Consequences, Source, and Suggestions. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8215–8228, Online. Association for Computational Linguistics.
  50. 50.Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books. arXiv:1506.06724 [cs].

Citation

MLA
Kirk, H., et al. “Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 1352–68, https://doi.org/10.18653/v1/2022.naacl-main.97.
APA
Kirk, H., Vidgen, B., Röttger, P., Thrush, T., & Hale, S. A. (2022). Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1352–1368. https://doi.org/10.18653/v1/2022.naacl-main.97
Chicago
Kirk, H., B. Vidgen, P. Röttger, T. Thrush, and S. A. Hale. 2022. “Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1352–68. https://doi.org/10.18653/v1/2022.naacl-main.97.
Harvard
Kirk, H. et al. (2022) “Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 1352–1368. Available at: https://doi.org/10.18653/v1/2022.naacl-main.97.
Vancouver
1. Kirk H, Vidgen B, Röttger P, Thrush T, Hale SA (2022) Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 1352–1368

BibTeX

@inproceedings{kirk-etal-2022-hatemoji,
    title = "{H}atemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate",
    author = "Kirk, Hannah  and
      Vidgen, Bertie  and
      Rottger, Paul  and
      Thrush, Tristan  and
      Hale, Scott A.",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.97/",
    doi = "10.18653/v1/2022.naacl-main.97",
    pages = "1352--1368"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/