Built independently by an author, for readers. Read the story and support ChapterPal

keyword

toxic language

Toxic language refers to any form of written or spoken communication that is disrespectful, abusive, hateful, or likely to demean, intimidate, and cause harm to individuals or groups. Commonly used as an umbrella concept in digital content moderation and natural language processing, it encompasses behaviors such as hate speech, cyberbullying, personal insults, violent threats, and discriminatory stereotyping. Because toxic language degrades constructive public discourse, fosters hostile environments, and causes emotional or psychological distress, research and automated safety systems focus on detecting, categorizing, and mitigating its occurrence across online platforms, conversational agents, and machine learning datasets.

2 items

The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics

The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics

Matthias Orlikowski, Paul Röttger, Philipp Cimiano, Dirk Hovy

OrganizationsBielefeld UniversityBocconi UniversityUniversity of Oxford

Why you should read this

Demonstrates that incorporating sociodemographic attributes into multi-annotator models fails to improve individual label prediction in toxic content detection, cautioning researchers against the ecological fallacy of reducing individual human perspectives to demographic group averages.

Many NLP tasks exhibit human label variation, where different annotators give different labels to the same texts. This variation is known to depend, at least in part, on the sociodemographics of annotators. Recent research aims to model individual annotator behaviour rather than predicting aggregated labels, and we would expect that sociodemographic information is useful for these models. On the other hand, the ecological fallacy states that aggregate group behaviour, such as the behaviour of the average female annotator, does not necessarily explain individual behaviour. To account for sociodemographics in models of individual annotator behaviour, we introduce group-specific layers to multi-annotator models. In a series of experiments for toxic content detection, we find that explicitly accounting for sociodemographic attributes in this way does not significantly improve model performance. This result shows that individual annotation behaviour depends on much more than just sociodemographics.

Added

2026-10-03

Ethical and social risks of harm from Language Models

Ethical and social risks of harm from Language Models

Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William Isaac, Sean Legassick, Geoffrey Irving, Iason Gabriel

OrganizationsCalifornia Institute of TechnologyGoogleUniversity College DublinUniversity of Toronto

Why you should read this

Systematically organizes the scattered landscape of language model risks into a structured taxonomy that makes potential harms identifiable and actionable for developers, researchers, and policymakers working to build safer AI systems.

This paper aims to help structure the risk landscape associated with large-scale Language Models (LMs). In order to foster advances in responsible innovation, an in-depth understanding of the potential risks posed by these models is needed. A wide range of established and anticipated risks are analysed in detail, drawing on multidisciplinary expertise and literature from computer science, linguistics, and social sciences. We outline six specific risk areas: I. Discrimination, Exclusion and Toxicity, II. Information Hazards, III. Misinformation Harms, V. Malicious Uses, V. Human-Computer Interaction Harms, VI. Automation, Access, and Environmental Harms. The first area concerns the perpetuation of stereotypes, unfair discrimination, exclusionary norms, toxic language, and lower performance by social group for LMs. The second focuses on risks from private data leaks or LMs correctly inferring sensitive information. The third addresses risks arising from poor, false or misleading information including in sensitive domains, and knock-on risks such as the erosion of trust in shared information. The fourth considers risks from actors who try to use LMs to cause harm. The fifth focuses on risks specific to LLMs used to underpin conversational agents that interact with human users, including unsafe use, manipulation or deception. The sixth discusses the risk of environmental harm, job automation, and other challenges that may have a disparate effect on different social groups or communities. In total, we review 21 risks in-depth. We discuss the points of origin of different risks and point to potential mitigation approaches. Lastly, we discuss organisational responsibilities in implementing mitigations, and the role of collaboration and participation. We highlight directions for further research, particularly on expanding the toolkit for assessing and evaluating the outlined risks in LMs.

Added

2026-02-21