On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research
Luiza PozzobonBeyza ErmisPatrick LewisSara Hooker
Reveals how unannounced updates to commercial toxicity evaluation APIs like Perspective alter benchmark leaderboards and invalidate prior research conclusions, while establishing concrete guidelines to ensure reproducible evaluations over time.
Toxicity evaluation is a vital component of safe artificial intelligence deployment, but human moderation at scale is costly and exposes evaluators to psychological harm. Consequently, researchers and developers rely heavily on automated commercial tools, such as the Perspective tool, to benchmark language models and assess toxicity mitigation techniques. However, commercial application programming interfaces (APIs) are frequently updated behind the scenes to fix flaws and biases, often without formal notifications or model versioning. This lack of transparency undermines scientific reproducibility and distorts comparative risk assessments across language models.
The article evaluates how unannounced updates to black-box toxicity detection APIs impact the reproducibility of published benchmarks and scientific conclusions over time. It demonstrates the extent to which scoring drift alters relative model rankings and distorts evaluations of newly proposed toxicity mitigation techniques.
To measure these effects, the authors rescored established text generation datasets and benchmarks using a recent version of the Perspective tool (evaluated in early 2023) and compared the results against their historical baselines. The scope of the evaluation included the 100,000-sentence RealToxicityPrompts dataset originally released in 2020, 37 commercial and open-source language models benchmarked under the Holistic Evaluation of Language Models (HELM) framework, and six prominent toxicity mitigation methods published between 2019 and 2023.
The findings show substantial shifts in toxicity measurements and model rankings. First, rescoring the RealToxicityPrompts dataset revealed a 49% reduction in the number of prompts classified as toxic and an overall 34% drop in average toxicity scores, with approximately 10,000 previously toxic prompts now classified as non-toxic. Second, rescoring model outputs in the HELM benchmark caused 24 rank changes across 13 models under the Toxic Fraction metric, shifting the position of some models by up to 12 places. Third, the perceived performance of toxicity mitigation techniques changed unevenly; for example, the Expected Maximum Toxicity of a recently published method dropped from 33.2% to 23.6%, shifting its relative standing against historical baselines.
These results demonstrate that comparing modern model outputs against published historical scores creates an invalid, non-standardized comparison. Reusing historical baseline scores can lead decision-makers and researchers to inaccurate conclusions regarding model safety, mitigation efficacy, and risk profiles. The common practice of scoring new model continuations while inheriting legacy prompt scores artificially distorts toxicity metrics and creates a false impression of system safety.
To establish reliable benchmarks, the article recommends concrete practices for researchers and API providers. Commercial API providers should systematically version their models and notify users of updates. Researchers must open-source their generated outputs, log exact scoring dates, and uniformly rescore all baseline generations when evaluating new techniques. Living benchmarks should implement a fixed control set of text prompts; if control scores change due to API updates, all benchmarked models must be rescored simultaneously.
The study's primary limitation is its dependence on publicly available model generations, which constrained the analysis to datasets and techniques with accessible text outputs from 2019 onward. While automated tools provide essential scalability, stakeholders should exercise caution when reviewing toxicity evaluations that mix scoring snapshots across different time periods.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). This paper establishes the widely adopted RealToxicityPrompts benchmark and standardizes evaluating neural toxic degeneration via commercial scoring APIs like Perspective.
- Paper: Time Waits for No One! Analysis and Challenges of Temporal Misalignment, Kelvin Luu et al. (2022). This study analyzes temporal misalignment and performance degradation in NLP systems across time, providing key foundational context for evaluating evolving language and updating APIs.
- Paper: Hatemoji: A Test Suite and Adversarially-Generated Dataset for Benchmarking and Detecting Emoji-Based Hate, Hannah Kirk et al. (2022). This work demonstrates critical blind spots and failure modes in commercial toxicity systems like Google Jigsaw's Perspective API under evolving linguistic inputs.
- Paper: BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation, Tianxiang Sun et al. (2022). This paper highlights how neural model-based evaluation metrics inherit biases and yield unreliable benchmark scores, motivating audits of automated evaluators.
- Paper: Language (Technology) is Power: A Critical Survey of “Bias” in NLP, Su Lin Blodgett et al. (2020). This survey provides essential conceptual framing regarding the mismeasurement and operationalization of harms and bias across NLP evaluation pipelines.
- Paper: Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Hakan Inan et al. (2023). This paper introduces Llama Guard as an open-weight, transparent safeguard model to overcome the limitations and opacity of commercial black-box moderation APIs.
- Paper: Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment, Vyas Raina et al. (2024). This study investigates the adversarial vulnerabilities and scoring instability of using language models as automated judges for benchmark evaluations.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). This paper establishes a multi-dimensional benchmarking framework for model trustworthiness that addresses broader evaluation and reliability challenges beyond single static APIs.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This survey systematically explores the opportunities, biases, and structural challenges of using automated LLM-based evaluators across diverse domains.
- Paper: Feedback Loops With Language Models Drive In-Context Reward Hacking, Alexander Pan et al. (2024). This work examines how optimizing against automated proxy metrics such as the Perspective API induces reward hacking and increases negative side effects over feedback loops.
