Built independently by an author, for readers. Read the story and support ChapterPal

keyword

toxicity detection

Toxicity detection is the automated process of identifying and classifying harmful, offensive, abusive, or inappropriate language within text or speech using natural language processing and machine learning algorithms. Applied widely in digital content moderation, social media platforms, and artificial intelligence safety, it aims to recognize and mitigate the spread of abusive behavior, hate speech, harassment, and other hostile discourse. Systems designed for toxicity detection evaluate both explicit expressions of harm, such as overt insults, threats, and profanity, and implicit forms of toxicity, such as coded prejudice, microaggressions, and subtle bias. These tools typically utilize trained classification models and deep neural networks to evaluate contextual nuance and semantic patterns, assigning scores or labels that quantify the likelihood and severity of toxic content.

1 item

Unveiling the Implicit Toxicity in Large Language Models

Unveiling the Implicit Toxicity in Large Language Models

Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, Minlie Huang

Why you should read this

Reveals that large language models can generate subtle, implicit toxicity that evades standard safety filters, and introduces a reinforcement learning attack framework that exposes these safety blind spots while providing training data to improve classifier defenses.

The open-endedness of large language models (LLMs) combined with their impressive capabilities may lead to new safety issues when being exploited for malicious use. While recent studies primarily focus on probing toxic outputs that can be easily detected with existing toxicity classifiers, we show that LLMs can generate diverse implicit toxic outputs that are exceptionally difficult to detect via simply zero-shot prompting. Moreover, we propose a reinforcement learning (RL) based attacking method to further induce the implicit toxicity in LLMs. Specifically, we optimize the language model with a reward that prefers implicit toxic outputs to explicit toxic and non-toxic ones. Experiments on five widely-adopted toxicity classifiers demonstrate that the attack success rate can be significantly improved through RL fine-tuning. For instance, the RL-finetuned LLaMA-13B model achieves an attack success rate of 90.04% on BAD and 62.85% on Davinci003. Our findings suggest that LLMs pose a significant threat in generating undetectable implicit toxic outputs. We further show that fine-tuning toxicity classifiers on the annotated examples from our attacking method can effectively enhance their ability to detect LLM-generated implicit toxic language. The code is publicly available at https://github.com/thu-coai/Implicit-Toxicity.

Added

2026-10-03