Hate Speech and Counter Speech Detection: Conversational Context Does Matter

Xinchen YuEduardo BlancoLingzi Hong

article2022NAACL69 citations

Presents a context-aware dataset of Reddit comments to demonstrate that incorporating conversational history substantially alters human annotations and significantly boosts neural network performance when detecting hate speech and counter speech.

Listen

Online hate speech continues to pose severe threats to individuals and society, often leading to real-world harm and targeted aggression. While platforms frequently combat this issue by blocking hateful users, an increasingly promising long-term strategy is counter speech, which directly challenges and neutralizes hateful rhetoric. Most existing automated content moderation systems evaluate comments in complete isolation. This approach risks misinterpreting user intent because online communication depends heavily on context. The article evaluates whether conversational context—specifically the immediately preceding comment—alters human perception of hate and counter speech, and demonstrates whether incorporating this context improves the accuracy of automated detection models.

To investigate these questions, the authors compiled a dataset of 6,846 English Reddit comment-reply pairs. Human crowd workers annotated these comments under two separate experimental setups: first by viewing the target reply alone, and second by viewing the target alongside its parent comment. Using this annotated benchmark, the authors trained and evaluated language models using a transformer-based neural network architecture. They tested context-unaware configurations against context-aware configurations that ingested both parent and reply texts, while also evaluating performance enhancements such as blending noisy annotations and pretraining on related conversational tasks like stance detection.

The findings establish that conversational context is essential for accurate classification. First, providing context fundamentally shifts the ground truth, causing human judgment to change for 38.3% of the evaluated comments; specifically, 34.2% of comments seen as hate in isolation and 55.1% seen as counter-hate in isolation shifted to neutral once context was visible. Second, context-aware machine learning models consistently outperformed isolated text models across all evaluation categories, achieving the highest overall performance score of 0.64 when combined with stance pretraining and data blending, compared to 0.58 for basic isolated models. Third, linguistic analysis revealed distinct contextual cues: counter speech frequently relies on question marks and problem-solving language, whereas hate speech concentrates high profanity directly within the reply. Finally, error analysis showed that incorporating context resolves major failure modes in automated systems, fixing 48% of errors caused by a lack of standalone information and 19% of errors driven by sarcasm or irony.

These findings have significant implications for platform governance, content moderation costs, and compliance risks. Conventional moderation tools that evaluate posts in isolation risk penalizing legitimate counter speech while failing to catch indirect or sarcastic hate speech. Inaccurate flagging can alienate users, infringe on open discourse, and create regulatory and safety liabilities. By demonstrating that conversation-level understanding significantly enhances detection precision, the article shows that digital safety systems must account for dialogue structure to remain reliable.

Organizations developing or deploying automated moderation tools should transition from single-comment classifiers to context-aware models that incorporate preceding conversational turns. Engineering teams should also leverage stance-detection data during model pretraining, as identifying agreement or disagreement substantially aids in distinguishing hate from counter-interventions. However, decision-makers should recognize existing limitations: the analysis was confined to Reddit discussions, utilized single-level parent context rather than full conversation threads, and relied on keyword-assisted data sampling. Further pilot testing across diverse platforms and expanded dialogue structures is recommended before broad operational rollout.

arXiv: 2206.06423
Cover for Hate Speech and Counter Speech Detection: Conversational Context Does Matter

Abstract

Hate speech is plaguing the cyberspace along with user-generated content. This paper investigates the role of conversational context in the annotation and detection of online hate and counter speech, where context is defined as the preceding comment in a conversation thread. We created a context-aware dataset for a 3-way classification task on Reddit comments: hate speech, counter speech, or neutral. Our analyses indicate that context is critical to identify hate and counter speech: human judgments change for most comments depending on whether we show annotators the context. A linguistic analysis draws insights into the language people use to express hate and counter speech. Experimental results show that neural networks obtain significantly better results if context is taken into account. We also present qualitative error analyses shedding light into (a) when and why context is beneficial and (b) the remaining errors made by our best model when context is taken into account.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Dataset Collection and Annotation
  • 3.1 Collecting (Parent, Target) pairs
  • 3.2 Annotation Guidelines
  • 4 Corpus Analysis
  • 5 Experiments and Results
  • 5.1 Quantitative Results
  • 6 Qualitative Analysis
  • 7 Conclusions and Future Work
  • 8 Ethical Considerations
  • Acknowledgements
  • References
  • A Annotation Interface
  • B Detailed Results
  • C Hyperparameters to Fine-tune the Systems

Knowls

  1. Knowl 1 — Three-Way Categorization of Conversational Online Speech

    definition

    In the context of conversational threads consisting of a preceding comment (denoted as ParentParent) and a responding comment (denoted as TargetTarget), user-generated target comments are categorized into three mutually exclusive classes:

    1. Hate: The author of the TargetTarget attacks an individual or a group with the explicit intention to vilify, humiliate, or incite hatred against them.
    2. Counter-hate (or Counter Speech): The author of the TargetTarget directly challenges, condemns, or calls out hateful content expressed in another comment (such as the ParentParent).
    3. Neutral: The author of the TargetTarget neither conveys hate nor opposes or challenges hate expressed in another comment.
  2. Knowl 2 — Contextual Reddit Hate and Counter-Speech Corpus

    data/table

    The contextual hate and counter-speech dataset contains 6,8466{,}846 English Reddit comment pairs (Parent,Target)(Parent, Target) extracted from 416416 discussion submissions using a lexicon of 1,7261{,}726 hate and harassment terms. Pairs where the same author wrote both ParentParent and TargetTarget are excluded. When taking the ParentParent context into account, the overall ground-truth label distribution is:

    • Neutral: 49%49\%
    • Hate: 28%28\%
    • Counter-hate: 23%23\%

    When the TargetTarget comment itself contains a hate keyword, the distribution shifts to 45%45\% Hate, 19%19\% Counter-hate, and 36%36\% Neutral.

    The dataset is partitioned into two quality subsets based on crowdsourced inter-annotator agreement using Krippendorff's alpha (α\alpha):

    • Gold Dataset: 4,7514{,}751 pairs where inter-annotator agreement reaches α≥0.6\alpha \ge 0.6. The Gold set is split into 70%70\% training, 15%15\% validation, and 15%15\% testing.
    • Silver Dataset: 2,0952{,}095 remaining lower-agreement pairs, utilized strictly as weak auxiliary supervision during model training.
  3. Knowl 3 — Iterative MACE-Based Annotator Filtering Algorithm

    algorithm

    To curate a high-quality Gold annotation subset from crowdsourced 5-fold redundant human annotations, annotators are iteratively pruned based on Multi-Annotator Competence Estimation (MACE) competence scores until inter-coder reliability satisfies Krippendorff's α≥0.6\alpha \ge 0.6.

    Input: Set of annotated items D, set of annotators A, initial annotation assignments
    Output: Adjudicated labels for Gold items D_gold, Silver items D_silver
    repeat
        Compute competence score C(a) for every annotator a in A using MACE
        Identify annotator a_min = argmin_{a in A} C(a)
        Remove all annotations submitted by a_min from D
        Remove a_min from A
        Calculate Krippendorff's alpha on the remaining annotations across D
    until alpha >= 0.6 or cannot remove further annotators
    Assign final ground truth labels using MACE competence-weighted adjudication:
    D_gold = {items in D with alpha >= 0.6}
    D_silver = {remaining items in original D not in D_gold}
    return D_gold, D_silver
  4. Knowl 4 — Impact of Conversational Context on Human Labeling of Hate and Counter Speech

    empirical result

    Showing the preceding context (ParentParent) causes human annotators to change their classification of the TargetTarget comment in 38.3%38.3\% of instances compared to evaluating the TargetTarget in isolation.

    Annotated Without Parent (Isolated Target)
    Annotated With Parent Hate (%) Counter-hate (%) Neutral (%)
    Hate 57.4 8.4 34.2
    Counter-hate 18.7 26.2 55.1
    Neutral 9.7 8.1 82.2

    Key shifts induced by conversational context include:

    • 55.1%55.1\% of comments judged as Counter-hate with context are deemed Neutral without context.
    • 34.2%34.2\% of comments judged as Hate with context appear Neutral in isolation.
    • 18.7%18.7\% of comments that are Counter-hate when considering context are misinterpreted as Hate in isolation due to the presence of vulgar or aggressive language targeted at the hate perpetrator.
    • 8.4%8.4\% of comments that are Hate with context are misinterpreted as Counter-hate in isolation.
  5. Knowl 5 — Linguistic Distinctions Between Hate and Counter Speech in Dialogues

    empirical result

    Statistical comparisons (unpaired tt-tests with Bonferroni correction for multiple hypothesis testing) between dialogue threads labeled as Counter-hate versus Hate reveal specific linguistic markers across the ParentParent and TargetTarget comments:

    1. Question Marks in Target: Significantly more frequent in Counter-hate TargetsTargets (p<0.001p < 0.001, passes Bonferroni correction), typically appearing in rhetorical questions challenging hateful assertions.
    2. Pronoun Distribution in Parent: ParentsParents that provoke Counter-hate contain fewer 1st person pronouns (II, meme; p<0.001p < 0.001, passes Bonferroni) and more 2nd person pronouns (youyou, youryour; p<0.001p < 0.001, passes Bonferroni), reflecting offensive targeting of others.
    3. Profanity Distribution: High profanity in the ParentParent comment strongly correlates with the TargetTarget being Counter-hate (p<0.001p < 0.001), whereas high profanity in the TargetTarget comment correlates with the TargetTarget being Hate (p<0.001p < 0.001).
    4. Cognitive and Problem-Solving Markers: Counter-hate TargetsTargets contain significantly more awareness words (p<0.001p < 0.001), enlightenment words (p<0.001p < 0.001), and problem-solving words (p<0.001p < 0.001).
    5. Valence and Disgust: ParentsParents provoking Counter-hate exhibit significantly higher negative words (p<0.001p < 0.001), whereas Counter-hate TargetsTargets contain fewer negative words (p<0.001p < 0.001) and fewer disgust words (p<0.001p < 0.001) compared to Hate TargetsTargets.
  6. Knowl 6 — Context-Aware Classification Architecture with Silver Data Blending and Stance Pretraining

    model/method

    The classification model is built upon a 12-layer pretrained RoBERTa-base transformer. To represent conversational context, the ParentParent and TargetTarget texts are concatenated with a separation token: [CLS] Parent [SEP] Target [SEP].

    The representation from the [CLS] token is passed through a fully connected layer (768768 units, tanh⁡\tanh activation) and a 3-unit linear layer with softmax activation over {Hate,Counter-hate,Neutral}\{\text{Hate}, \text{Counter-hate}, \text{Neutral}\}.

    Two training enhancement strategies are combined:

    1. Silver Data Blending: Training runs for mm blending epochs followed by nn Gold-only epochs. In blending epoch k∈{1,…,m}k \in \{1, \dots, m\}, the network is fed all Gold samples and a fraction fkf_k of Silver samples, where the fraction is decayed by a factor α∈[0,1]\alpha \in [0, 1] each epoch (f1=1.0f_1 = 1.0, fk=max⁡(0,1−(k−1)α)f_k = \max(0, 1 - (k-1)\alpha)). For context-aware modeling without auxiliary pretraining, optimal α=0.3\alpha = 0.3; when combined with stance pretraining, optimal α=1.0\alpha = 1.0.
    2. Auxiliary Task Pretraining: Sequential pretraining on conversational stance detection (agree, neutral, attack) using the DEBAGREEMENT corpus prior to fine-tuning on the target dataset.
  7. Knowl 7 — Classification Performance of Context-Aware vs. Isolated Models

    empirical result

    Evaluating models on the 15%15\% Gold test split shows that incorporating ParentParent context consistently improves F1F_1-scores across all categories, with auxiliary stance pretraining and Silver blending yielding the highest performance.

    Hate Counter-hate Neutral Weighted Average
    Model Setup P R F1 P R F1 P R F1 P R F1
    Majority Baseline 0.00 0.00 0.00 0.00 0.00 0.00 0.51 1.00 0.67 0.26 0.51 0.34
    Trained with Target 0.56 0.55 0.56 0.41 0.36 0.38 0.67 0.71 0.69 0.58 0.59 0.58
    + Silver 0.58 0.55 0.57 0.44 0.42 0.43 0.69 0.72 0.70 0.60 0.61 0.61
    + Stance Pretrain 0.56 0.55 0.56 0.51 0.41 0.45 0.68 0.74 0.71 0.61 0.61 0.61
    + Silver + Stance 0.55 0.56 0.56 0.49 0.53 0.51 0.67 0.69 0.70 0.61 0.61 0.61
    Trained with Parent_Target 0.56 0.62 0.59 0.52 0.38 0.44 0.68 0.72 0.70 0.61 0.62 0.61
    + Silver 0.58 0.57 0.57 0.49 0.51 0.50 0.72 0.71 0.72 0.63 0.63 0.63
    + Stance Pretrain 0.55 0.66 0.60 0.54 0.43 0.48 0.71 0.70 0.71 0.63 0.63 0.63
    + Silver + Stance 0.55 0.65 0.60 0.54 0.52 0.53 0.74 0.68 0.71 0.64 0.64 0.64

    The context-aware model combining Silver data and stance pretraining significantly outperforms the context-unaware baseline (p<0.01p < 0.01, McNemar's test), achieving a weighted average F1F_1 of 0.640.64 compared to 0.580.58 for the context-free model.

  8. Knowl 8 — Error Taxonomy: Context-Resolved Errors and Residual Failure Modes

    empirical result

    Qualitative analysis reveals the primary error classes resolved by context awareness as well as the remaining failure modes of the best context-aware model (Parent_Target + Silver + Stance):

    Errors Fixed by Context Awareness (TargetTarget-only errors resolved by Parent_TargetParent\_Target):

    1. Lack of information (48%48\%): TargetTarget meaning is underspecified or referential (e.g., answering a rhetorical question), requiring ParentParent to identify hateful intent.
    2. Negation (27%27\%): TargetTarget scolds or counters hateful premises via negation, mistaken for Neutral in isolation.
    3. Sarcasm or irony (19%19\%): Sarcastic remarks that attack hateful arguments, misclassified as Hate without context.
    4. Hate without swear words (8%8\%): Non-profane statements that introduce or validate hatred only relative to the ParentParent.

    Remaining Error Modes in the Best Context-Aware Model:

    1. Negation and Double Negation (28%28\%): Multi-clause or double negation confuses the model's stance attribution (e.g., 'Don't forget male isn't a gender, it's a disease.' predicted as Counter-hate instead of Hate).
    2. Rhetorical Questions (27%27\%): Aggressive or disdainful rhetorical questions directed at the ParentParent author misclassified as Counter-hate instead of Hate.
    3. Swear Word Discrepancy (16%16\% total): Non-hateful comments containing profanity misclassified as Hate (8%8\%), and hateful comments lacking profanity misclassified as Counter-hate or Neutral (8%8\%).
    4. Intricate / Nuanced Text (7%7\%): Expressing strong negative attitudes or agreements without vilifying groups, mistakenly labeled as Hate.
  9. Knowl 9 — Limitations of 1-Turn Context and Lexical Seed Sampling

    limitation

    The methodology exhibits three main limitations:

    1. Context Window Constraint: Context is strictly restricted to the immediate preceding comment (ParentParent). While multi-turn conversational trees could provide deeper context, they increase the risk of biasing human annotators' personal stance.
    2. Lexicon Sampling Bias: Reddit conversation retrieval relies on an initial lexicon of 1,7261{,}726 hate terms, which biases the collected samples toward explicit lexical triggers, partially mitigated by extracting conversational replies to keyword comments.
    3. MACE Filtering Pruning Disagreements: Pruning low-competence annotators to reach α≥0.6\alpha \ge 0.6 results in assigning contentious or ambiguous boundary cases to the Silver training partition, preventing controversial cases from appearing in the evaluation test set.

Coverage note — None was omitted. All principal contributions—including task definitions, dataset curation and MACE filtering, linguistic SEANCE analysis, neural context modeling with Silver blending/stance pretraining, experimental quantitative results, and qualitative error taxonomy—are covered.

References

  1. 1.Ron Artstein and Massimo Poesio. 2008. Inter-coder agreement for computational linguistics. Comput. Linguist., 34(4):555–596.
  2. 2.Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. The pushshift reddit dataset. In Proceedings of the Fourteenth International AAAI Conference on Web and Social Media, ICWSM 2020, Held Virtually, Original Venue: Atlanta, Georgia, USA, June 8-11, 2020, pages 830–839. AAAI Press.
  3. 3.Manuela Caiani, Benedetta Carlotti, and Enrico Padoan. 2021. Online hate speech and the radical right in times of pandemic: The italian and english cases. Javnost - The Public, 28(2):202–218.
  4. 4.Yi-Ling Chung, Elizaveta Kuzmenko, Serra Sinem Tekiroglu, and Marco Guerini. 2019. CONAN - COunter NArratives through nichesourcing: a multilingual dataset of responses to fight online hate speech. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2819–2829, Florence, Italy. Association for Computational Linguistics.
  5. 5.Scott A Crossley, Kristopher Kyle, and Danielle S McNamara. 2017. Sentiment analysis and social cognition engine (seance): An automatic tool for sentiment, social cognition, and social-order analysis. Behavior research methods, 49(3):803–821.
  6. 6.Thomas Davidson, Dana Warmsley, Michael W. Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. In Proceedings of the Eleventh International Conference on Web and Social Media, ICWSM 2017, Montréal, Québec, Canada, May 15-18, 2017, pages 512–515. AAAI Press.
  7. 7.Subhabrata Dutta, Dipankar Das, and Tanmoy Chakraborty. 2020. Changing views: Persuasion modeling and argument extraction from online discussions. Information Processing & Management, 57(2):102085.
  8. 8.Mai ElSherief, Vivek Kulkarni, Dana Nguyen, William Yang Wang, and Elizabeth M. Belding. 2018. Hate lingo: A target-based linguistic analysis of hate speech in social media. In Proceedings of the Twelfth International Conference on Web and Social Media, ICWSM 2018, Stanford, California, USA, June 25-28, 2018, pages 42–51. AAAI Press.
  9. 9.Margherita Fanton, Helena Bonaldi, Serra Sinem Tekiroglu, and Marco Guerini. 2021. Human-in-the-loop for data collection: a multi-target counter narrative dataset to fight online hate speech. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3226–3240, Online. Association for Computational Linguistics.
  10. 10.Tracie Farrell, Miriam Fernandez, Jakub Novotny, and Harith Alani. 2019. Exploring misogyny across the manosphere in reddit. In Proceedings of the 10th ACM Conference on Web Science, WebSci ’19, page 87–96, New York, NY, USA. Association for Computing Machinery.
  11. 11.Paula Fortuna and Sérgio Nunes. 2018. A survey on automatic detection of hate speech in text. ACM Comput. Surv., 51(4).
  12. 12.Lei Gao and Ruihong Huang. 2017. Detecting online hate speech using context aware models. In Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017, pages 260–266, Varna, Bulgaria. INCOMA Ltd.
  13. 13.Joshua Garland, Keyan Ghazi-Zahedi, Jean-Gabriel Young, Laurent Hébert-Dufresne, and Mirta Galesic. 2020. Countering hate on social media: Large scale classification of hate and counter speech. In Proceedings of the Fourth Workshop on Online Abuse and Harms, pages 102–112, Online. Association for Computational Linguistics.
  14. 14.Debanjan Ghosh, Avijit Vajpayee, and Smaranda Muresan. 2020. A report on the 2020 sarcasm detection shared task. In Proceedings of the Second Workshop on Figurative Language Processing, pages 1–11, Online. Association for Computational Linguistics.
  15. 15.Bing He, Caleb Ziems, Sandeep Soni, Naren Ramakrishnan, Diyi Yang, and Srijan Kumar. 2021. Racism is a virus: Anti-asian hate and counterspeech in social media during the covid-19 crisis. In Proceedings of the 2021 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining, page 90–94, New York, NY, USA. Association for Computing Machinery.
  16. 16.Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard Hovy. 2013. Learning whom to trust with MACE. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1120–1130, Atlanta, Georgia. Association for Computational Linguistics.
  17. 17.Nicola F Johnson, R Leahy, N Johnson Restrepo, Nicolas Velasquez, Ming Zheng, P Manrique, P Devkota, and Stefan Wuchty. 2019. Hidden resilience and adaptive dynamics of the global online hate ecology. Nature, 573(7773):261–265.
  18. 18.Klaus Krippendorff. 2011. Computing krippendorff’s alpha-reliability. https://repository.upenn.edu/asc_papers/43/. Accessed: 2021-02-08.
  19. 19.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  20. 20.Binny Mathew, Anurag Illendula, Punyajoy Saha, Soumya Sarkar, Pawan Goyal, and Animesh Mukherjee. 2020. Hate begets hate: A temporal study of hate speech. Proceedings of the ACM on Human-Computer Interaction, 4(CSCW2):1–24.
  21. 21.Binny Mathew, Punyajoy Saha, Hardik Tharad, Subham Rajgaria, Prajwal Singhania, Suman Kalyan Maity, Pawan Goyal, and Animesh Mukherjee. 2019. Thou shalt not hate: Countering online hate speech. In Proceedings of the Thirteenth International Conference on Web and Social Media, ICWSM 2019, Munich, Germany, June 11-14, 2019, pages 369–380. AAAI Press.
  22. 22.Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. Hatexplain: A benchmark dataset for explainable hate speech detection. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 14867–14875. AAAI Press.
  23. 23.Quinn McNemar. 1947. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2):153–157.
  24. 24.Stefano Menini, Alessio Palmero Aprosio, and Sara Tonelli. 2021. Abuse is contextual, what about nlp? the role of context in abusive language annotation and detection. CoRR, abs/2103.14916.
  25. 25.Chikashi Nobata, Joel R. Tetreault, Achint Thomas, Yashar Mehdad, and Yi Chang. 2016. Abusive language detection in online user content. In Proceedings of the 25th International Conference on World Wide Web, WWW 2016, Montreal, Canada, April 11 - 15, 2016, pages 145–153. ACM.
  26. 26.Alexandra Olteanu, Carlos Castillo, Jeremy Boy, and Kush R. Varshney. 2018. The effect of extremist violence on hateful speech online. In Proceedings of the Twelfth International Conference on Web and Social Media, ICWSM 2018, Stanford, California, USA, June 25-28, 2018, pages 221–230. AAAI Press.
  27. 27.John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain, and Ion Androutsopoulos. 2020. Toxicity detection: Does context really matter? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4296–4305, Online. Association for Computational Linguistics.
  28. 28.John Pougué-Biyong, Valentina Semenova, Alexandre Matton, Rachel Han, Aerin Kim, Renaud Lambiotte, and Doyne Farmer. 2021. DEBAGREEMENT: A comment-reply dataset for (dis)agreement detection in online debates. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  29. 29.Rob Procter, Helena Webb, Pete Burnap, William Housley, Adam Edwards, Matthew L. Williams, and Marina Jirotka. 2019. A study of cyber hate on twitter with implications for social media governance strategies. In Proceedings of the 2019 Truth and Trust Online Conference (TTO 2019), London, UK, October 4-5, 2019.
  30. 30.Yada Pruksachatkun, Phil Yeres, Haokun Liu, Jason Phang, Phu Mon Htut, Alex Wang, Ian Tenney, and Samuel R. Bowman. 2020. jiant: A software toolkit for research on general-purpose text understanding models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 109–117, Online. Association for Computational Linguistics.
  31. 31.Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding, and William Yang Wang. 2019. A benchmark dataset for learning to intervene in online hate speech. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4755–4764, Hong Kong, China. Association for Computational Linguistics.
  32. 32.Yafeng Ren, Yue Zhang, Meishan Zhang, and Donghong Ji. 2016. Context-sensitive twitter sentiment classification using neural network. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, AAAI’16, page 215–221. AAAI Press.
  33. 33.Mohammadreza Rezvan, Saeedeh Shekarpour, Lakshika Balasuriya, Krishnaprasad Thirunarayan, Valerie L. Shalin, and Amit Sheth. 2018. A quality type-aware annotated corpus and lexicon for harassment research. In Proceedings of the 10th ACM Conference on Web Science, WebSci ’18, page 33–36, New York, NY, USA. Association for Computing Machinery.
  34. 34.Robert D Richards and Clay Calvert. 2000. Counterspeech 2000: A new look at the old remedy for bad speech. BYU L. Rev., page 553.
  35. 35.Sara Rosenthal, Noura Farra, and Preslav Nakov. 2017. SemEval-2017 task 4: Sentiment analysis in Twitter. In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017), pages 502–518, Vancouver, Canada. Association for Computational Linguistics.
  36. 36.Marta Sabou, Kalina Bontcheva, Leon Derczynski, and Arno Scharl. 2014. Corpus annotation through crowdsourcing: Towards best practice guidelines. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14), pages 859–866, Reykjavik, Iceland. European Language Resources Association (ELRA).
  37. 37.Anna Schmidt and Michael Wiegand. 2017. A survey on hate speech detection using natural language processing. In Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, pages 1–10, Valencia, Spain. Association for Computational Linguistics.
  38. 38.Eyal Shnarch, Carlos Alzate, Lena Dankin, Martin Gleize, Yufang Hou, Leshem Choshen, Ranit Aharonov, and Noam Slonim. 2018. Will it blend? blending weak and strong labeled data in a neural network for argumentation mining. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 599–605, Melbourne, Australia. Association for Computational Linguistics.
  39. 39.Bertie Vidgen, Dong Nguyen, Helen Margetts, Patricia Rossini, and Rebekah Tromble. 2021. Introducing CAD: the contextual abuse dataset. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2289–2303, Online. Association for Computational Linguistics.
  40. 40.Kenneth D Ward. 1997. Free speech and the development of liberal virtues: An examination of the controversies involving flag-burning and hate speech. University of Miami Law Review, 52(3):733–792.
  41. 41.Zeerak Waseem and Dirk Hovy. 2016. Hateful symbols or hateful people? predictive features for hate speech detection on Twitter. In Proceedings of the NAACL Student Research Workshop, pages 88–93, San Diego, California. Association for Computational Linguistics.
  42. 42.Ziqi Zhang and Lei Luo. 2019. Hate speech detection: A solved problem? the challenging case of long tail on twitter. Semantic Web, 10(5):925–945.
  43. 43.Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik, Kalina Bontcheva, Trevor Cohn, and Isabelle Augenstein. 2018. Discourse-aware rumour stance classification in social media using sequential classifiers. Information Processing & Management, 54(2):273–290.

Citation

MLA
Yu, X., et al. “Hate Speech and Counter Speech Detection: Conversational Context Does Matter”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 5918–30, https://doi.org/10.18653/v1/2022.naacl-main.433.
APA
Yu, X., Blanco, E., & Hong, L. (2022). Hate Speech and Counter Speech Detection: Conversational Context Does Matter. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5918–5930. https://doi.org/10.18653/v1/2022.naacl-main.433
Chicago
Yu, X., E. Blanco, and L. Hong. 2022. “Hate Speech and Counter Speech Detection: Conversational Context Does Matter”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 5918–30. https://doi.org/10.18653/v1/2022.naacl-main.433.
Harvard
Yu, X., Blanco, E. and Hong, L. (2022) “Hate Speech and Counter Speech Detection: Conversational Context Does Matter”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 5918–5930. Available at: https://doi.org/10.18653/v1/2022.naacl-main.433.
Vancouver
1. Yu X, Blanco E, Hong L (2022) Hate Speech and Counter Speech Detection: Conversational Context Does Matter. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 5918–5930

BibTeX

@inproceedings{yu-etal-2022-hate,
    title = "Hate Speech and Counter Speech Detection: Conversational Context Does Matter",
    author = "Yu, Xinchen  and
      Blanco, Eduardo  and
      Hong, Lingzi",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.433/",
    doi = "10.18653/v1/2022.naacl-main.433",
    pages = "5918--5930"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/