keyword
content moderation
Content moderation is the practice of monitoring, reviewing, and filtering user-generated digital content to ensure it complies with established platform rules, legal standards, and community guidelines. This process aims to identify and address inappropriate, toxic, or harmful material such as hate speech, harassment, misinformation, and explicit content to maintain safe and constructive online environments. Moderation is conducted using human reviewers, automated machine learning and natural language processing systems, or hybrid approaches that combine algorithmic detection with human evaluation. Because determining what constitutes offensive or inappropriate material often involves subjective judgment, effective moderation requires navigating challenges related to cultural nuances, differing annotator perspectives, and the balance between automated efficiency and contextual understanding.
2 items

When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
Eve Fleisig, Rediet Abebe, Dan Klein
Why you should read this
Presents a framework that predicts individual annotator judgments alongside the targeted demographic groups of text to identify when targeted populations disagree with majority-vote labels in offensive content detection.
People often disagree on subjective tasks such as determining what is offensive or toxic online, where each annotator brings their own perspective influenced by factors like culture, identity, and lived experience. In many annotation settings today, however, we ask multiple people — who may have different beliefs — to provide just one label per example, treating majority vote as ground truth. We show how training models to capture individual annotator behavior instead yields better modeling of disagreement patterns among human raters.
Added
2026-10-02

ChatGPT outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, Maël Kubli
Why you should read this
Demonstrates that zero-shot ChatGPT surpasses crowd-workers in both accuracy and intercoder reliability across multiple text-annotation tasks while reducing annotation costs by twenty-fold.
Many NLP applications require manual data annotations for a variety of tasks, notably to train classifiers or evaluate the performance of unsupervised models. Depending on the size and degree of complexity, the tasks may be conducted by crowd-workers on platforms such as MTurk as well as trained annotators, such as research assistants. Using a sample of 2,382 tweets, we demonstrate that ChatGPT outperforms crowd-workers for several annotation tasks, including relevance, stance, topics, and frames detection. Specifically, the zero-shot accuracy of ChatGPT exceeds that of crowd-workers for four out of five tasks, while ChatGPT's intercoder agreement exceeds that of both crowd-workers and trained annotators for all tasks. Moreover, the per-annotation cost of ChatGPT is less than $0.003 -- about twenty times cheaper than MTurk. These results show the potential of large language models to drastically increase the efficiency of text classification.
Added
2026-09-24
