Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter

Zeerak WaseemDirk Hovy

article2016NAACL1,893 citationsACL 2026 Test of Time Award

Presents a critical race theory-based annotation framework and a benchmark dataset of 16,000 tweets to evaluate how demographic and extra-linguistic user features compare against character n-grams for automated hate speech detection.

Listen

Online hate speech represents an escalating challenge for social media platforms, posing direct risks to user safety due to its established connection to real-world hate crimes. Platform moderation currently relies heavily on manual human review, which is difficult to scale, introduces subjective individual bias, and struggles to balance content removal with free speech protections. The article addresses the urgent need for standardized definitions and automated methods to identify hate speech accurately, specifically examining whether incorporating extra-linguistic user information improves automated detection.

The article evaluated the effectiveness of character-level textual features alongside demographic and geographic data for detecting hate speech on Twitter. Specifically, it set out to determine whether identifying user demographics, such as gender and geographic location, improves identification performance over text-based models alone.

To conduct this evaluation, the researchers established an 11-point annotation guide grounded in critical race theory to standardize what constitutes hate speech. Using these criteria, they collected and manually labeled an unbalanced dataset of 16,914 English-language tweets gathered over two months, categorizing them into sexist, racist, or non-offensive content. They used logistic regression classifiers with ten-fold cross-validation to compare various feature combinations, assessing character sequences (character n-grams) against word-level tokens, text lengths, geographic location, and inferred user gender.

The investigation yielded several critical findings regarding feature effectiveness. First, models using character sequences significantly outperformed traditional word-based models, improving the F1 performance score from 64.58 to 73.89. Second, adding extra-linguistic features generally degraded classification performance: incorporating user location reduced the F1 score to 73.62, and incorporating text length reduced it to 73.66. Third, while adding inferred gender produced the highest overall F1 score of 73.93, this minor improvement was not statistically significant. Fourth, the demographic analysis revealed a strong gender imbalance among identified perpetrators, with men accounting for 50.24% of sexist tweets and 33.33% of racist tweets, while women accounted for under 2.3% in each category.

These findings suggest that automated moderation systems should prioritize sub-word character analysis over demographic profiling and metadata. Because abusive language frequently repurposes ordinary words or introduces spelling variations, character-level modeling provides greater accuracy and resilience against sparse text data. Furthermore, relying on user metadata introduces noise that can harm system performance while adding unnecessary computational complexity and data governance burdens.

For organizations developing automated content moderation pipelines, the immediate recommendation is to deploy character-based sequence modeling as the baseline text analysis tool rather than relying solely on static keyword blocklists or demographic metadata. Operational teams should also establish formal, structured annotation guidelines, such as the criteria introduced in the article, to reduce annotator subjectivity during human review. Before demographic signals can be reliably used in production systems, further work is required to develop more comprehensive user profile inference methods.

Decision-makers should view these results with caution regarding demographic factors due to major data coverage limitations. Gender could only be identified for 52.56% of users in the dataset, and location metadata was similarly scarce. Additionally, the racist subset within the dataset was concentrated among a very small pool of users, which may limit generalizability across wider social media environments.

  • Paper: Automated Hate Speech Detection and the Problem of Offensive Language, Thomas Davidson et al. (2017). This paper directly builds upon and refines the source's hate speech detection paradigm by separating general offensive language from targeted hate speech to prevent widespread false positives.
  • Paper: Beyond Accuracy: Behavioral Testing of NLP Models with CheckList, Marco Túlio Ribeiro et al. (2020). This work introduces behavioral and linguistic capability testing for NLP models, offering a systematic framework to evaluate error modes and demographic biases that emerge in classifiers like those in the source.
  • Paper: A Survey on Bias and Fairness in Machine Learning, Ninareh Mehrabi et al. (2019). This comprehensive survey synthesizes downstream algorithmic bias and fairness concerns across machine learning applications, generalizing issues encountered in demographic-sensitive text classification.
Cover for Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter

Abstract

Hate speech in the form of racist and sexist remarks are a common occurrence on social media. For that reason, many social media services address the problem of identifying hate speech, but the definition of hate speech varies markedly and is largely a manual effort (BBC, 2015; Lomas, 2015).

We provide a list of criteria founded in critical race theory, and use them to annotate a publicly available corpus of more than 16k tweets. We analyze the impact of various extra-linguistic features in conjunction with character n-grams for hate-speech detection. We also present a dictionary based the most indicative words in our data.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Data
  • 3 Demographic distribution
  • 4 Lexical distribution
  • 5 Geographic distribution
  • 6 Evaluation
  • 7 Related Work
  • 8 Conclusion
  • References

Knowls

  1. Knowl 1 — Annotation Criteria for Hate Speech Identification

    definition

    An 11-point decision list founded in critical race theory (derived by negating structural privileges and identifying mechanisms used to silence or oppress minorities) defines whether a tweet is offensive hate speech:

    A tweet is offensive if it:

    1. Uses a sexist or racial slur.
    2. Attacks a minority.
    3. Seeks to silence a minority.
    4. Criticizes a minority without a well-founded argument.
    5. Promotes, but does not directly use, hate speech or violent crime.
    6. Criticizes a minority using a straw man argument.
    7. Blatantly misrepresents truth or seeks to distort views on a minority with unfounded claims.
    8. Shows support of problematic hashtags (e.g., #BanIslam, #whoriental, #whitegenocide), where problematic hashtags are defined as terms fulfilling one or more of the other criteria.
    9. Negatively stereotypes a minority.
    10. Defends xenophobia or sexism.
    11. Contains an offensive screen name according to the previous criteria, where the tweet content is ambiguous at best, and the tweet addresses a topic that satisfies any of the above criteria.
  2. Knowl 2 — Impact of Extra-Linguistic and Linguistic Features on Hate Speech Detection

    empirical result

    A logistic regression classifier evaluated via 10-fold cross-validation on a 16,914-tweet dataset evaluates the predictive utility of character nn-grams compared to word nn-grams and extra-linguistic metadata (gender, location, length):

    Metric char n-grams +gender +gender +loc word n-grams
    F1 73.89 73.93 73.62* 64.58
    Precision 72.87% 72.93% 72.58% 64.39%
    Recall 77.75% 77.74% 77.43% 71.93%

    *Statistically significant degradation (p=0.0355p = 0.0355) tested via a 10,000-sample bootstrap test at p<0.05p < 0.05.

    Additional feature configurations:

    • Character nn-grams (1--4) + length features (tweet length, user description length, average word length): F1=73.66\text{F}_1 = 73.66.
    • Character nn-grams (1--4) + gender + location + length: F1=73.47\text{F}_1 = 73.47.

    Character nn-grams of lengths 1 to 4 provide the strongest textual representation, outperforming word nn-grams by over 9 points in F1. Gender is the only extra-linguistic feature that yields a slight positive gain over character nn-grams alone (from 73.89 to 73.93 F1), though the difference is not statistically significant. Adding geographic location markers or text length features degrades performance.

  3. Knowl 3 — Twitter Hate Speech Dataset Composition and Collection Procedure

    data/table

    A corpus of English tweets was collected via the public Twitter search API over a two-month period using bootstrapped queries consisting of identity terms, common slurs, and topical hashtags (e.g., MKR, asian drive, feminazi, immigrant, nigger, sjw, WomenAgainstFeminism, blameonenotall, islam terrorism, notallmen, victimcard, victim card, arab terror, gamergate, jsil, racecard, race card). From 136,052 retrieved tweets, 16,914 were manually annotated without class balancing to reflect real-world distributions.

    The class distribution across tweets and author accounts is:

    • Sexist content: 3,383 tweets sent by 613 distinct users.
    • Racist content: 1,972 tweets sent by 9 distinct users (demonstrating a highly concentrated set of prolific authors).
    • Neither sexist nor racist: 11,559 tweets sent by 614 distinct users.

    Total annotated dataset size: 16,914 tweets.

  4. Knowl 4 — Inter-Annotator Agreement in Hate Speech Annotation

    empirical result

    Annotation of the 16,914 tweets by the authors and an independent external reviewer (a female gender studies student) produced an inter-annotator agreement of Cohen's κ=0.84\kappa = 0.84.

    Disagreements were heavily concentrated in the sexism category, which accounted for 85% of all annotation divergences. Furthermore, 98% of all reviewer reclassifications shifted an initial hate speech label to "neither sexist nor racist" (with the remaining 2% changed to racist). Disagreements primarily arose from differing thresholds for missing context: the initial annotators classified tweets targeting individual female participants of reality television shows (e.g., My Kitchen Rules) as sexist, whereas the reviewer required explicit gendered animus.

  5. Knowl 5 — User Demographic Gender Distribution in Hate Speech Data

    data/table

    User demographic gender is extracted by matching user profile descriptions, screen names, and given names against male and female name dictionaries alongside gendered pronouns, honorifics, and gendered nouns. This method identified gender for 52.56% of users across the dataset.

    The distribution of inferred gender across tweet classes is:

    Gender All Racism Sexism Neither
    Men 50.08% 33.33% 50.24% 50.92%
    Women 2.26% 0.00% 2.28% 1.74%
    Unidentified 47.64% 66.66% 47.47% 47.32%

    Identified perpetrators are overwhelmingly male across all classes. However, because the inferred gender distribution for sexist tweets (50.24% men, 2.28% women) is nearly identical to that for benign tweets (50.92% men, 1.74% women), gender features provide minimal discriminative utility for classification.

  6. Knowl 6 — Indicative Lexical and Sub-Word Features for Sexism and Racism

    empirical result

    Model coefficients summed across 10-fold cross-validation in logistic regression identify distinct feature signatures distinguishing racist and sexist tweets:

    • Racist tweets are characterized by standard religious and ethnic terminology re-appropriated into hostile discourse. The most frequent full words are islam (1.44%), muslims (1.01%), muslim (0.65%), not (0.53%), mohammed (0.52%), religion (0.40%), isis (0.38%), jews (0.37%), prophet (0.36%), and #islam (0.35%). The most indicative character nn-gram coefficients are 'sl', 'sla', 'slam', 'isla', 'l', 'a', 'isl', 'lam', 'i', 'e', 'mu', 's', 'am', 'm', 'la', 'is', 'slim', 'musl', 'usli', and 'lim'.
    • Sexist tweets contain explicit gendered slurs, female references, and topic markers related to television broadcasts. The most frequent full words are not (1.83%), sexist (1.68%), #mkr (1.57%), women (0.83%), kat (0.57%), girls (0.48%), like (0.42%), call (0.36%), #notsexist (0.36%), and female (0.34%). The most indicative character nn-gram coefficients are 'xist', 'sexi', 'ka', 'sex', 'kat', 'exis', 'xis', 'exi', 'xi', 'bitc', 'ist', 'bit', 'itch', 'itc', 'fem', 'ex', 'bi', 'irl', 'wom', and 'girl'.
  7. Knowl 7 — Tweet Character Length Distribution Across Classes

    data/table

    Tweet lengths measured in characters (excluding whitespace) vary systematically across hate speech categories:

    Statistic Racism Sexism None
    Mean 60.47 52.93 47.95
    Std. 17.44 21.16 23.43
    Min. 11.00 2.00 2.00
    Max. 115.00 118.00 129.00

    Racist tweets exhibit the highest mean character length (60.47 characters) with the lowest standard deviation (17.44), whereas non-hate tweets are shorter on average (47.95 characters) with wider variance. Despite these distributional differences, incorporating length statistics as classification features degrades model F1 score.

Coverage note — None was omitted; all primary contributions—including the critical race theory annotation guidelines, dataset construction details, inter-annotator statistics, demographic extraction and distributions, character vs. word n-gram benchmarking, feature ablations, and indicative n-gram analyses—are represented.

References

  1. 1.Diana Abbas. 2015. What’s in a location. https://www.youtube.com/watch?v=GNlDO9Lt8J8, October. Talk at Twitter Flight 2015. Seen on Jan 17th 2016.
  2. 2.BBC. 2015. Facebook, google and twitter agree german hate speech deal. http://www.bbc.com/news/world-europe-35105003. Accessed on 26/11/2016.
  3. 3.Ying Chen, Yilu Zhou, Sencun Zhu, and Heng Xu. 2012a. Detecting offensive language in social media to protect adolescent online safety. In Privacy, Security, Risk and Trust (PASSAT), 2012 International Conference on and 2012 International Conference on Social Computing (SocialCom), pages 71–80. IEEE, September.
  4. 4.Yunfei Chen, Lanbo Zhang, Aaron Michelony, and Yi Zhang. 2012b. 4is of social bully filtering: Identity, inference, influence, and intervention. In Proceedings of the 21st ACM International Conference on Information and Knowledge Management, CIKM ’12, pages 2677–2679, New York, NY, USA. ACM.
  5. 5.Tori DeAngelis. 2009. Unmasking ’racial micro aggressions’. Monitor on Psychology, 40(2):42.
  6. 6.Lisa Eadicicco. 2014. This female game developer was harassed so severely on twitter she had to leave her home. http://www.businessinsider.com/brianna-wu-harassed-twitter-2014-10?IR=T, Oct. Seen on Jan. 25th, 2016.
  7. 7.Stephan Gouws, Donald Metzler, Congxing Cai, and Eduard Hovy. 2011. Contextual bearing on linguistic variation in social media. In Proceedings of the Workshop on Languages in Social Media, LSM ’11, pages 20–29, Stroudsburg, PA, USA. Association for Computational Linguistics.
  8. 8.Mark Kantrowitz. 1994. Name corpus: List of male, female, and pet names. http://www.cs.cmu.edu/afs/cs/project/ai-repository/ai/areas/nlp/corpora/names/0.html. Last accessed on 29th February 2016.
  9. 9.Heather Hensman Kettrey and Whitney Nicole Laster. 2014. Staking territory in the world white web: An exploration of the roles of overt and color-blind racism in maintaining racial boundaries on a popular web site. Social Currents, 1(3):257–274.
  10. 10.Natasha Lomas. 2015. Facebook, google, twitter commit to hate speech action in germany. http://techcrunch.com/2015/12/16/germany-fights-hate-speech-on-social-media/, Dec. Seen on 23rd Jan. 2016.
  11. 11.Alexandra Ma. 2015. Global survey finds nordic countries have the most feminists. http://www.huffingtonpost.com/entry/global-gender-equality-study-yougov_us_564604cce4b045bf3deeb96d, November. Seen on Jan 19th.
  12. 12.Black Lives Matter. 2012. Guiding principles. http://blacklivesmatter.com/guiding-principles/. Accessed on 26/11/2016.
  13. 13.Peggy McIntosh, 2003. Understanding prejudice and discrimination., chapter White privilege: Unpacking the invisible knapsack, pages 191–196. McGraw-Hill.
  14. 14.Geir Moulson. 2016. Zuckerberg in germany: No place for hate speech on facebook. http://abcnews.go.com/Technology/wireStory/zuckerberg-place-hate-speech-facebook-37217309. Accessed 10/03/2016.
  15. 15.Colin Roberts, Martin Innes, Matthew Williams, Jasmin Tregidga, and David Gadd. 2013. Understanding who commits hate crime and why they do it.
  16. 16.Sara Sood, Judd Antin, and Elizabeth Churchill. 2012a. Profanity use in online communities. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, pages 1481–1490. ACM.
  17. 17.Sara Owsley Sood, Judd Antin, and Elizabeth F. Churchill. 2012b. Using crowdsourcing to improve profanity detection. In AAAI Spring Symposium: Wisdom of the Crowd, volume SS-12-06 of AAAI Technical Report. AAAI.
  18. 18.Stephan Tulkens, Lisa Hilte, Elise Lodewyckx, Ben Verhoeven, and Walter Daelemans. 2015. Detecting racism in dutch social media posts, 2015/12/18.
  19. 19.William Warner and Julia Hirschberg. 2012. Detecting hate speech on the world wide web. In Proceedings of the Second Workshop on Language in Social Media, LSM ’12, pages 19–26, Stroudsburg, PA, USA. Association for Computational Linguistics.
  20. 20.Hate Speech Watch. 2014. Hate crimes: Consequences of hate speech. http://www.nohatespeechmovement.org/hate-speech-watch/focus/consequences-of-hate-speech, June. Seen on on 23rd Jan. 2016.
  21. 21.Julia Carrie Wong. 2016. Mark Zuckerberg tells Facebook staff to stop defacing Black Lives Matter slogans. http://www.theguardian.com/technology/2016/feb/25/mark-zuckerberg-facebook-defacing-black-lives-matter-signs. Accessed on 10/03/2016.
  22. 22.Matthew Zook. 2012. Mapping racist tweets in response to president obama’s re-election. http://www.floatingsheep.org/2012/11/mapping-racist-tweets-in-response-to.html. Accessed on 11/03/2016.

Citation

MLA
Waseem, Z., and D. Hovy. “Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter”. Proceedings of the NAACL Student Research Workshop, 2016, pp. 88–93, https://doi.org/10.18653/v1/N16-2013.
APA
Waseem, Z., & Hovy, D. (2016). Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter. Proceedings of the NAACL Student Research Workshop, 88–93. https://doi.org/10.18653/v1/N16-2013
Chicago
Waseem, Z., and D. Hovy. 2016. “Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter”. Proceedings of the NAACL Student Research Workshop, 88–93. https://doi.org/10.18653/v1/N16-2013.
Harvard
Waseem, Z. and Hovy, D. (2016) “Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter”, Proceedings of the NAACL Student Research Workshop. Association for Computational Linguistics, pp. 88–93. Available at: https://doi.org/10.18653/v1/N16-2013.
Vancouver
1. Waseem Z, Hovy D (2016) Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter. In: Proceedings of the NAACL Student Research Workshop. Association for Computational Linguistics, pp 88–93

BibTeX

@inproceedings{Waseem_2016, title={Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter}, url={http://dx.doi.org/10.18653/v1/N16-2013}, DOI={10.18653/v1/n16-2013}, booktitle={Proceedings of the NAACL Student Research Workshop}, publisher={Association for Computational Linguistics}, author={Waseem, Zeerak and Hovy, Dirk}, year={2016}, pages={88–93} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF