On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research

Luiza PozzobonBeyza ErmisPatrick LewisSara Hooker

article2023EMNLP61 citations

Reveals how unannounced updates to commercial toxicity evaluation APIs like Perspective alter benchmark leaderboards and invalidate prior research conclusions, while establishing concrete guidelines to ensure reproducible evaluations over time.

Listen

Toxicity evaluation is a vital component of safe artificial intelligence deployment, but human moderation at scale is costly and exposes evaluators to psychological harm. Consequently, researchers and developers rely heavily on automated commercial tools, such as the Perspective tool, to benchmark language models and assess toxicity mitigation techniques. However, commercial application programming interfaces (APIs) are frequently updated behind the scenes to fix flaws and biases, often without formal notifications or model versioning. This lack of transparency undermines scientific reproducibility and distorts comparative risk assessments across language models.

The article evaluates how unannounced updates to black-box toxicity detection APIs impact the reproducibility of published benchmarks and scientific conclusions over time. It demonstrates the extent to which scoring drift alters relative model rankings and distorts evaluations of newly proposed toxicity mitigation techniques.

To measure these effects, the authors rescored established text generation datasets and benchmarks using a recent version of the Perspective tool (evaluated in early 2023) and compared the results against their historical baselines. The scope of the evaluation included the 100,000-sentence RealToxicityPrompts dataset originally released in 2020, 37 commercial and open-source language models benchmarked under the Holistic Evaluation of Language Models (HELM) framework, and six prominent toxicity mitigation methods published between 2019 and 2023.

The findings show substantial shifts in toxicity measurements and model rankings. First, rescoring the RealToxicityPrompts dataset revealed a 49% reduction in the number of prompts classified as toxic and an overall 34% drop in average toxicity scores, with approximately 10,000 previously toxic prompts now classified as non-toxic. Second, rescoring model outputs in the HELM benchmark caused 24 rank changes across 13 models under the Toxic Fraction metric, shifting the position of some models by up to 12 places. Third, the perceived performance of toxicity mitigation techniques changed unevenly; for example, the Expected Maximum Toxicity of a recently published method dropped from 33.2% to 23.6%, shifting its relative standing against historical baselines.

These results demonstrate that comparing modern model outputs against published historical scores creates an invalid, non-standardized comparison. Reusing historical baseline scores can lead decision-makers and researchers to inaccurate conclusions regarding model safety, mitigation efficacy, and risk profiles. The common practice of scoring new model continuations while inheriting legacy prompt scores artificially distorts toxicity metrics and creates a false impression of system safety.

To establish reliable benchmarks, the article recommends concrete practices for researchers and API providers. Commercial API providers should systematically version their models and notify users of updates. Researchers must open-source their generated outputs, log exact scoring dates, and uniformly rescore all baseline generations when evaluating new techniques. Living benchmarks should implement a fixed control set of text prompts; if control scores change due to API updates, all benchmarked models must be rescored simultaneously.

The study's primary limitation is its dependence on publicly available model generations, which constrained the analysis to datasets and techniques with accessible text outputs from 2019 onward. While automated tools provide essential scalability, stakeholders should exercise caution when reviewing toxicity evaluations that mix scoring snapshots across different time periods.

arXiv: 2304.12397
Cover for On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research

Abstract

Perception of toxicity evolves over time and often differs between geographies and cultural backgrounds. Similarly, black-box commercially available APIs for detecting toxicity, such as the Perspective API, are not static, but frequently retrained to address any unattended weaknesses and biases. We evaluate the implications of these changes on the reproducibility of findings that compare the relative merits of models and methods that aim to curb toxicity. Our findings suggest that research that relied on inherited automatic toxicity scores to compare models and techniques may have resulted in inaccurate findings. Rescoring all models from HELM, a widely respected living benchmark, for toxicity with the recent version of the API led to a different ranking of widely used foundation models. We suggest caution in applying apples-to-apples comparisons between studies and lay recommendations for a more structured approach to evaluating toxicity over time. 1

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 RealToxicityPrompts (RTP)
  • 2.2 Evaluating Model Toxicity
  • 3 Results
  • 3.1 REALToxicityPrompts Distribution Changes
  • 3.2 Impact of API Changes on Rankings of Model Risk
  • 3.2.1 Impact on Living Benchmarks
  • 3.3 Impact of API Changes on Reproducibility of Research Contributions
  • 4 Recommendations
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • References
  • A Perspective API
  • B Qualitative evaluation of RTP prompts
  • C UDDIA Rescored Results
  • D HELM
  • D.1 Toxicity Measurement Details

Knowls

  1. Knowl 1 — Perspective API rescoring substantially changed the REALTOXICITYPROMPTS dataset

    empirical result

    The authors rescored the same REALTOXICITYPROMPTS (RTP) texts with the Perspective API in February 2023, while the released dataset had been scored before September 2020. A prompt was classified as toxic when its Perspective toxicity score was greater than 0.5. The published-versus-rescored statistics shown on page 4 report 21,744 versus 11,676 toxic prompts and 77,272 versus 87,475 non-toxic prompts, respectively. Thus, the number of toxic prompts fell by 49% under the newer API scores. The authors also report a 34% reduction in average toxicity scores and approximately 10,000 transitions from toxic to non-toxic, compared with only 232 transitions from non-toxic to toxic.

    The joint prompt-and-continuation score distribution also changed markedly: the rescored dataset contained 48,600 sequences (49%) in the [0, 0.25) toxicity range, 25,796 (26%) in [0.25, 0.5), 19,719 (20%) in [0.5, 0.75), and 5,228 (5%) in [0.75, 1.0]. The released dataset had been constructed with approximately 25,000 sequences in each range. The largest shift was away from the highest-toxicity bin, indicating that many texts considered toxic under the earlier API standards were no longer classified as toxic by the newer version.

  2. Knowl 2 — Three-way rescoring protocol for measuring API-version effects

    experimental setup

    The study evaluates toxicity under three scoring conditions while keeping the underlying prompts and generated texts fixed: (1) the originally published prompt and generation scores, (2) newly rescored generations combined with inherited prompt scores, and (3) newly rescored prompts and generations scored under the same Perspective API version. The authors treat the second condition as an invalid apples-to-oranges comparison because prompt and continuation scores reflect incompatible toxicity definitions; the third condition is the technically consistent measurement.

    For a language model, the evaluation protocol conditions the model on a prompt, generates a continuation, and scores the generated text with the Perspective API. Expected Maximum Toxicity (EMT) is the maximum toxicity score among k=25k=25 generations for one prompt. Toxicity Probability (TP) is the empirical probability that at least one of those 25 generations has toxicity greater than 0.5. Toxic Fraction (TF) is the fraction of generated instances whose toxicity score is greater than 0.5. Results are reported separately for toxic and non-toxic prompts.

  3. Knowl 3 — Mixing old prompt scores with new generation scores biases toxicity estimates

    empirical result

    The three-way evaluation of GPT1, GPT2, GPT3, CTRL, and CTRL-W shows that changing only the generation scores produces lower EMT and TP values for both toxic and non-toxic prompts relative to the originally published results. When the prompt scores are also rescored, the toxicity metrics rise again, particularly for toxic prompts. Consequently, retaining inherited prompt labels while rescoring only model continuations can make a model appear less toxic than it would under a single, consistent Perspective API version.

    This result applies even though the texts being evaluated are identical in all three conditions. The apparent change in model toxicity therefore arises from the interaction between API-version-dependent prompt classification and API-version-dependent continuation scores, rather than from a change in the language model or its generated text.

  4. Knowl 4 — Perspective API rescoring changed HELM model rankings

    empirical result

    The authors rescored the REALTOXICITYPROMPTS generations in HELM v0.2.2 with the Perspective API in April 2023 and compared the results with the static benchmark scores. HELM contained 37 language models, each evaluated on the same 1,000 toxic or non-toxic prompts. For Toxic Fraction, 13 models changed metric values and those changes produced 24 ranking changes. For Expected Maximum Toxicity, 18 models changed metric values and those changes produced 21 ranking changes. The average absolute change across models was 0.018 for Toxic Fraction and 0.041 for Expected Maximum Toxicity.

    The largest Toxic Fraction change was for openai_text-curie-001, whose score decreased from 0.107 to 0.090 and whose rank improved from 35th to 23rd. Its Expected Maximum Toxicity decreased by 10.8%, moving it from 34th to 23rd. The page-3 bump plots show substantial rank crossings for both Toxic Fraction and Expected Maximum Toxicity. Models already near the least-toxic end were comparatively stable, with Cohere models consistently remaining among the ten least-toxic models.

  5. Knowl 5 — Replicated toxicity-mitigation results changed under newer API scores

    empirical result

    The authors rescored open-sourced generations from six toxicity-mitigation or baseline systems: DAPT, DExperts (Large), GPT2 (Large), GeDi, PPLM (10%), and UDDIA with threshold TH=40. The evaluations used a selection of 10,000 non-toxic prompts according to the published prompt scores. The comparison was possible because the relevant generations had been released publicly.

    For UDDIA, a result published only a few months before this study, Expected Maximum Toxicity decreased from 33.2% to 23.6% after rescoring the same generations. The page-7 plots show corresponding changes in Toxic Fraction and Toxicity Probability. Min-max-normalized comparisons showed that the direction and magnitude of change were not uniform across systems: most baseline models had nearly zero variation, whereas UDDIA and DAPT had non-zero and inconsistent slopes. Relative to the other systems, UDDIA was perceived as more toxic under the newer scoring behavior, while DAPT was perceived as less toxic. Therefore, toxicity-mitigation rankings and conclusions about mitigation–quality trade-offs can change over time even when the models and generations remain fixed.

  6. Knowl 6 — API drift affected multiple Perspective attributes, not only toxicity

    empirical result

    The authors compared published and February 2023 Perspective API score distributions for the RTP prompts across several attributes. Distributional changes were measured with Wasserstein distance, interpreted as the minimum amount of work needed to transform one score distribution into the other. The largest changes were for threat, with distance 0.189, and severe toxicity, with distance 0.153. The distances were 0.097 for toxicity, 0.093 for profanity, and 0.046 for flirtation. Thus, the observed instability is not confined to the toxicity attribute used for the main experiments; updates also altered the API's outputs for threat, severe toxicity, profanity, flirtation, and other moderation-related attributes.

  7. Knowl 7 — Score changes disproportionately affected identity-related and non-American text

    empirical result

    The qualitative analysis of RTP prompts whose labels changed between the released and rescored versions found asymmetric patterns. The toxic-to-non-toxic shift affected approximately 10,000 prompts and was prominent among texts containing words such as “black,” “sex,” and “white.” The much smaller non-toxic-to-toxic group contained about 200 prompts and often included slurs or strongly abusive wording.

    The largest individual score changes also exposed possible contextual sensitivity. Examples of toxic-to-non-toxic transitions included an Indonesian football-club description changing from 0.87 to 0.01, an Internet-of-Things product prompt changing from 0.78 to 0.02, and a passage about arthropods changing from 0.84 to 0.11. Examples in the opposite direction included profanity-containing prompts changing from 0.11 to 0.89 and from 0.21 to 0.91. The authors characterize these patterns as evidence that API updates can alter the treatment of non-American names, regions, and technical vocabulary as well as explicitly abusive language.

  8. Knowl 8 — Recommended version-controlled protocol for toxicity evaluation

    model/method

    The paper recommends a reproducibility protocol for research and living benchmarks that depend on black-box toxicity APIs:

    • API maintainers should version the underlying scoring models and consistently notify users whenever a model update changes scores.
    • Authors should release model generations, their toxicity scores, evaluation code when possible, and the date on which every text was scored.
    • When comparing a new toxicity-mitigation method with prior work, authors should rescore all available open-sourced generations. They should treat generations that cannot be rescored as having potentially outdated scores rather than as directly comparable evidence.
    • A living benchmark should maintain a fixed control set of sequences and rescore that set whenever a new model is added. If the control-set metrics change, the benchmark should rescore earlier models. If an earlier model cannot be rescored because its outputs are inaccessible, the benchmark should explicitly mark the result as outdated or remove it from that benchmark version.

    The recommendations are intended to prevent incompatible API snapshots from being combined into one comparison and to make the scoring date and version part of the reported experimental provenance.

  9. Knowl 9 — HELM toxicity measurements differ from the standard RTP protocol

    experimental setup

    The authors found that HELM v0.2.2 does not implement the standard REALTOXICITYPROMPTS evaluation exactly. HELM generates five continuations per prompt rather than the standard RTP value of 25. Its default generation requests use top-p=1 and top-k=1; when a model's API gives precedence to top-k, the five continuations may be identical or nearly identical. HELM computes Toxic Fraction over unique text instances, so five textually identical toxic completions count as one toxic text while the denominator remains five instances.

    HELM also uses the first spanScore returned by the Perspective API, whereas RTP uses the summaryScore. The summaryScore covers the whole sequence, while a spanScore covers a relevant section detected by the API. These implementation differences mean that HELM toxicity values are not necessarily directly comparable with results from studies that closely follow the standard RTP protocol, independently of API-version drift.

  10. Knowl 10 — Scope limitation from dependence on open-sourced generations

    limitation

    The replication analysis is limited to toxicity-mitigation studies whose model continuations were released publicly or whose authors provided access. It therefore cannot assess studies with inaccessible generations, even if those studies relied on the same black-box toxicity API. The benchmark sample focuses on toxicity-mitigation techniques proposed or published from 2019 through 2023; work published before 2019 was outside the reported scope and could only be added if its continuations were available for rescoring.

Coverage note — The complete 37-model HELM metric matrix and the full ten-example prompt table were omitted as lower-level repetitions; their load-bearing aggregate changes and representative score transitions are retained.

References

  1. 1.Reuben Binns, Michael Veale, Max Van Kleek, and Nigel Shadbolt. 2017. Like trainer, like bot? inheritance of bias in algorithmic content moderation. In Social Informatics: 9th International Conference, SocInfo 2017, Oxford, UK, September 13-15, 2017, Proceedings, Part II 9, pages 405–415. Springer.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Lingjiao Chen, Matei Zaharia, and James Zou. 2023. How is chatgpt’s behavior changing over time? arXiv preprint arXiv:2307.09009.
  4. 4.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  5. 5.Jon F Claerbout and Martin Karrenbach. 1992. Electronic documents give reproducible research a new meaning. In SEG technical program expanded abstracts 1992, pages 601–604. Society of Exploration Geophysicists.
  6. 6.K Bretonnel Cohen, Jingbo Xia, Pierre Zweigenbaum, Tiffany Callahan, Orin Hargraves, Foster Goss, Nancy Ide, Aurélie Névéol, Cyril Grouin, and Lawrence Hunter. 2018. Three dimensions of reproducibility in natural language processing. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018).
  7. 7.Brandon Dang, Martin J Riedl, and Matthew Lease. 2018. But who protects the moderators? the case of crowdsourced image moderation. arXiv preprint arXiv:1804.10999.
  8. 8.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019. Plug and play language models: A simple approach to controlled text generation. arXiv preprint arXiv:1912.02164.
  9. 9.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  10. 10.Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019. Racial bias in hate speech and abusive language detection datasets. arXiv preprint arXiv:1905.12516.
  11. 11.Farshid Faal, Ketra Schmitt, and Jia Yuan Yu. 2022. Reward modeling for mitigating toxicity in transformer-based language models. Applied Intelligence, pages 1–15.
  12. 12.SK Gargee, Pranav Bhargav Gopinath, Shridhar Reddy SR Kancharla, CR Anand, and Anoop S Babu. 2022. Analyzing and addressing the difference in toxicity prediction between different comments with same semantic meaning in google’s perspective api. In ICT Systems and Sustainability: Proceedings of ICT4SD 2022, pages 455–464. Springer.
  13. 13.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462.
  14. 14.Aaron Gokaslan and Vanya Cohen. 2019. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus.
  15. 15.Nitesh Goyal, Ian D Kivlichan, Rachel Rosen, and Lucy Vasserman. 2022. Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW2):1–28.
  16. 16.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online. Association for Computational Linguistics.
  17. 17.Shagun Jhaver, Iris Birman, Eric Gilbert, and Amy Bruckman. 2019. Human-machine collaboration for content regulation: The case of reddit automoderator. ACM Transactions on Computer-Human Interaction (TOCHI), 26(5):1–35.
  18. 18.Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858.
  19. 19.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2020. Gedi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367.
  20. 20.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2021. GeDi: Generative discriminator guided sequence generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4929–4952, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  21. 21.Alyssa Lees, Vinh Q Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman. 2022. A new generation of perspective api: Efficient multilingual character-level transformers. arXiv preprint arXiv:2202.11176.
  22. 22.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  23. 23.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A Smith, and Yejin Choi. 2021a. Dexperts: Decoding-time controlled text generation with experts and anti-experts. arXiv preprint arXiv:2105.03023.
  24. 24.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021b. DExperts: Decoding-time controlled text generation with experts and anti-experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6691–6706, Online. Association for Computational Linguistics.
  25. 25.Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency, pages 220–229.
  26. 26.Roger D Peng. 2011. Reproducible research in computational science. Science, 334(6060):1226–1227.
  27. 27.Hans E Plesser. 2018. Reproducibility vs. replicability: a brief history of a confused terminology. Frontiers in neuroinformatics, 11:76.
  28. 28.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018. Improving language understanding by generative pre-training.
  29. 29.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
  30. 30.Laura Ruis, Akbir Khan, Stella Biderman, Sara Hooker, Tim Rocktäschel, and Edward Grefenstette. 2022. Large language models are not zero-shot communicators. arXiv preprint arXiv:2210.14986.
  31. 31.Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A Smith. 2019. The risk of racial bias in hate speech detection. In Proceedings of the 57th annual meeting of the association for computational linguistics, pages 1668–1678.
  32. 32.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  33. 33.Shoaib Ahmed Siddiqui, Nitarshan Rajkumar, Tegan Maharaj, David Krueger, and Sara Hooker. 2022. Metadata archaeology: Unearthing data subsets by leveraging training dynamics.
  34. 34.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990.
  35. 35.Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The psychological well-being of content moderators: the emotional labor of commercial moderation and avenues for improving support. In Proceedings of the 2021 CHI conference on human factors in computing systems, pages 1–14.
  36. 36.Rachael Tatman, Jake VanderPlas, and Sohier Dane. 2018. A practical taxonomy of reproducibility for machine learning research.
  37. 37.Michael Veale and Reuben Binns. 2017. Fairer machine learning in the real world: Mitigating discrimination without collecting sensitive data. Big Data & Society, 4:205395171774353.
  38. 38.Boxin Wang, Wei Ping, Chaowei Xiao, Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Bo Li, Anima Anandkumar, and Bryan Catanzaro. 2022. Exploring the limits of domain-adaptive training for detoxifying large-scale language models. arXiv preprint arXiv:2202.04173.
  39. 39.Zeerak Waseem. 2016. Are you a racist or am i seeing things? annotator influence on hate speech detection on twitter. In Proceedings of the first workshop on NLP and computational social science, pages 138–142.
  40. 40.Johannes Welbl, Amelia Glaese, Jonathan Uesato, Sumanth Dathathri, John Mellor, Lisa Anne Hendricks, Kirsty Anderson, Pushmeet Kohli, Ben Coppin, and Po-Sen Huang. 2021. Challenges in detoxifying language models. arXiv preprint arXiv:2109.07445.
  41. 41.Zonghan Yang, Xiaoyuan Yi, Peng Li, Yang Liu, and Xing Xie. 2022. Unified detoxifying and debiasing in language generation via inference-time adaptive optimization. arXiv preprint arXiv:2210.04492.
  42. 42.Donglin Zhuang, Xingyao Zhang, Shuaiwen Song, and Sara Hooker. 2022. Randomness in neural network training: Characterizing the impact of tooling. In Proceedings of Machine Learning and Systems, volume 4, pages 316–336.

Citation

MLA
Pozzobon, L., et al. “On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 7595–609, https://doi.org/10.18653/v1/2023.emnlp-main.472.
APA
Pozzobon, L., Ermis, B., Lewis, P., & Hooker, S. (2023). On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7595–7609. https://doi.org/10.18653/v1/2023.emnlp-main.472
Chicago
Pozzobon, L., B. Ermis, P. Lewis, and S. Hooker. 2023. “On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7595–7609. https://doi.org/10.18653/v1/2023.emnlp-main.472.
Harvard
Pozzobon, L. et al. (2023) “On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 7595–7609. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.472.
Vancouver
1. Pozzobon L, Ermis B, Lewis P, Hooker S (2023) On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 7595–7609

BibTeX

@inproceedings{pozzobon-etal-2023-challenges,
    title = "On the Challenges of Using Black-Box {API}s for Toxicity Evaluation in Research",
    author = "Pozzobon, Luiza  and
      Ermis, Beyza  and
      Lewis, Patrick  and
      Hooker, Sara",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.472/",
    doi = "10.18653/v1/2023.emnlp-main.472",
    pages = "7595--7609"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/