Built independently by an author, for readers. Read the story and support ChapterPal

keyword

toxicity evaluation

Toxicity evaluation refers to the systematic assessment and measurement of harmful, offensive, abusive, or socially unacceptable text within natural language processing systems and artificial intelligence models. This process determines the likelihood and severity with which a model generates or propagates toxic content, such as hate speech, personal attacks, profanity, harassment, and identity-based bias, often in response to standard or adversarial prompts. In machine learning safety research, toxicity evaluation is conducted using dedicated benchmark datasets, automated classification tools, third-party scoring interfaces, or human annotation frameworks to assign quantitative toxicity scores. Because interpretations of unacceptable language vary across cultural contexts and linguistic norms, these evaluations serve as a critical mechanism for identifying safety vulnerabilities, benchmarking foundation models, and assessing the effectiveness of alignment and mitigation techniques.

1 item

On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research

On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research

Luiza Pozzobon, Beyza Ermis, Patrick Lewis, Sara Hooker

OrganizationsCohereRecod.aiSchool of Electrical and Computer EngineeringUniversity of Campinas

Why you should read this

Reveals how unannounced updates to commercial toxicity evaluation APIs like Perspective alter benchmark leaderboards and invalidate prior research conclusions, while establishing concrete guidelines to ensure reproducible evaluations over time.

Perception of toxicity evolves over time and often differs between geographies and cultural backgrounds. Similarly, black-box commercially available APIs for detecting toxicity, such as the Perspective API, are not static, but frequently retrained to address any unattended weaknesses and biases. We evaluate the implications of these changes on the reproducibility of findings that compare the relative merits of models and methods that aim to curb toxicity. Our findings suggest that research that relied on inherited automatic toxicity scores to compare models and techniques may have resulted in inaccurate findings. Rescoring all models from HELM, a widely respected living benchmark, for toxicity with the recent version of the API led to a different ranking of widely used foundation models. We suggest caution in applying apples-to-apples comparisons between studies and lay recommendations for a more structured approach to evaluating toxicity over time. 1

Added

2026-09-26