keyword
toxicity metrics
Toxicity metrics are quantitative evaluation measures used in natural language processing and content moderation to assess the degree to which text contains harmful, abusive, offensive, or disrespectful language. These metrics typically rely on automated machine learning classifiers, scoring algorithms, or specialized evaluation tools that analyze generated or human-authored text to assign numerical probabilities or categorical classifications indicating the presence and severity of toxic content. Commonly applied in artificial intelligence safety benchmarking and online platform governance, toxicity metrics help researchers and developers evaluate the effectiveness of alignment techniques, track harmful outputs across model iterations, and detect specific categories of undesirable content such as hate speech, harassment, insults, and profanity.
1 item

