Unveiling the Implicit Toxicity in Large Language Models

Jiaxin WenPei KeHao SunZhexin ZhangChengfei LiJinfeng BaiMinlie Huang

article2023EMNLP73 citations

Reveals that large language models can generate subtle, implicit toxicity that evades standard safety filters, and introduces a reinforcement learning attack framework that exposes these safety blind spots while providing training data to improve classifier defenses.

Listen

As large language models become widely deployed across consumer and enterprise applications, automated safety filters are essential to prevent harmful or abusive outputs. While existing safety mechanisms reliably catch explicit toxicity containing profanity or overt hate speech, bad actors can exploit generative models to convey toxicity through subtle, indirect language. Addressing implicit toxicity has become an urgent priority, as deployed systems risk spreading harmful content undetected if existing moderation tools fail to recognize nuanced harmful rhetoric.

The article demonstrates that large language models possess an innate ability to generate implicit toxic content that consistently bypasses current state-of-the-art moderation systems. It also establishes a reinforcement learning framework that further optimizes language models to produce highly evasive implicit toxicity, while evaluating whether retraining safety classifiers on this generated data can close the resulting security gap.

The researchers evaluated their methods on a standard benchmark dialogue dataset using a three-stage machine learning pipeline. First, they warm-started a base model using automated zero-shot examples generated by an instruction-tuned model. Second, they trained a reward model to score and prefer implicitly toxic responses over overtly toxic or benign ones, while incorporating penalties from an existing moderation filter. Third, they fine-tuned open-source language models—scaling from 1.3 billion to 13 billion parameters—using reinforcement learning to optimize against this reward function. The evasiveness of the generated text was then evaluated across five widely used industry and research toxicity classifiers, with ground truth verified by independent human annotators.

The findings reveal that standard commercial and open-source toxicity classifiers are highly vulnerable to large language model outputs. Zero-shot prompting alone generated implicit toxicity that bypassed classifiers 58% to nearly 97% of the time. Reinforcement learning fine-tuning exacerbated these vulnerabilities significantly: on a fine-tuned 13-billion-parameter model, the rate of undetected toxic responses reached 90% against a standard dialogue classifier and roughly 63% against an advanced language-model-based evaluator. Additionally, larger language models proved substantially more effective at generating evasive content because they combine multiple complex rhetorical devices, such as sarcasm, euphemism, and circumlocution. Despite these vulnerabilities, fine-tuning moderation classifiers on just 4,000 annotated examples of these evasive responses improved their detection rates against implicit toxicity from under 38% to roughly 79% without compromising general accuracy.

These results demonstrate a critical operational and compliance risk: current automated moderation systems provide a false sense of security, allowing subtly harmful outputs to slip through undetected. Organizations deploying generative AI risk reputational damage and policy violations if they rely solely on conventional toxicity filters. However, the study also proves that the same adversarial techniques can be used proactively to red-team models and strengthen automated moderation layers prior to deployment.

Organizations developing and deploying generative AI systems should implement red-teaming pipelines that specifically test for implicit toxicity rather than relying solely on explicit keyword detection. Moderation classifiers should be retrained and augmented using datasets composed of subtle, multi-feature toxic examples. Furthermore, because detecting implicit toxicity often requires contextual knowledge and linguistic reasoning, organizations should explore multi-layered evaluation systems that prompt evaluator models to explicitly analyze semantic nuances before classifying text as safe.

The primary limitation of this work involves noise in the automatically annotated comparison data used during reward modeling, which occasionally mislabeled subtle toxicity as benign. Additionally, the experimental scope was limited to models up to 13 billion parameters due to compute constraints, meaning the evasive potential of frontier-scale models remains an area requiring ongoing investigation. Nevertheless, the findings provide high confidence that current moderation classifiers suffer from systemic vulnerabilities against implicit toxicity and demonstrate a viable pathway to remediate these weaknesses.

Cover for Unveiling the Implicit Toxicity in Large Language Models

Abstract

The open-endedness of large language models (LLMs) combined with their impressive capabilities may lead to new safety issues when being exploited for malicious use. While recent studies primarily focus on probing toxic outputs that can be easily detected with existing toxicity classifiers, we show that LLMs can generate diverse implicit toxic outputs that are exceptionally difficult to detect via simply zero-shot prompting. Moreover, we propose a reinforcement learning (RL) based attacking method to further induce the implicit toxicity in LLMs. Specifically, we optimize the language model with a reward that prefers implicit toxic outputs to explicit toxic and non-toxic ones. Experiments on five widely-adopted toxicity classifiers demonstrate that the attack success rate can be significantly improved through RL fine-tuning. For instance, the RL-finetuned LLaMA-13B model achieves an attack success rate of 90.04% on BAD and 62.85% on Davinci003. Our findings suggest that LLMs pose a significant threat in generating undetectable implicit toxic outputs. We further show that fine-tuning toxicity classifiers on the annotated examples from our attacking method can effectively enhance their ability to detect LLM-generated implicit toxic language. The code is publicly available at https://github.com/thu-coai/Implicit-Toxicity.

Table of Contents

  • 1 Introduction
  • 2 Preliminary Experiments on Implicit Toxicity in Large Language Models
  • 3 Methodology
  • 3.1 Supervised Learning
  • 3.2 Reward Model Training
  • 3.3 Reinforcement Learning
  • 4 Experiments
  • 4.1 Settings
  • 4.2 Baselines
  • 4.3 Attacked Toxicity Classifiers
  • 4.4 Metrics
  • 4.5 Main Results
  • 4.6 Analysis
  • 4.7 Analysis of Implicit Toxicity in LLMs
  • 4.8 Improving Toxicity Classifiers
  • 5 Related Work
  • 5.1 Safety Issues of Language Models
  • 5.2 Toxicity Detection
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Proximal Policy Optimization
  • B Implementation Details
  • B.1 Data Preprocessing
  • B.2 Training Details
  • B.3 Example Prompts
  • B.3.1 Supervised Learning Data Collection
  • B.3.2 Comparison Data Collection
  • B.3.3 Toxicity Classification
  • C Generated Examples
  • D Instructions for Human Annotation
  • E Scaling Properties of Linguistic Features

Knowls

  1. Knowl 1 — A three-stage reinforcement-learning attack induces implicit toxicity

    model/method

    The method trains a language model to produce toxic responses that imply harmful content without relying on overtly offensive wording. It has three stages. First, an instruction-tuned model such as GPT-3.5-turbo generates responses to queries; the prompts request non-toxic, explicitly toxic, and implicitly toxic responses. A trainable policy is warm-started on the generated query-response pairs with maximum-likelihood training. Second, the current policy generates candidate responses for each query. GPT-3.5-turbo labels each response as implicit toxic, explicit toxic, or non-toxic; implicit-toxic responses are preferred to the other two classes, which are treated as equally preferred. The resulting comparisons train a reward model with the pairwise loss −log⁡σ(Rθ(x,yw)−Rθ(x,yl))-\log\sigma(R_\theta(x,y^w)-R_\theta(x,y^l)), where xx is a query, ywy^w and yly^l are the preferred and less-preferred responses, RθR_\theta is a scalar reward model, and σ\sigma is the logistic function. Third, proximal policy optimization (PPO) updates the policy using a reward that combines the learned preference signal, a toxicity-classifier penalty, and a KL penalty against the supervised policy:

    R^θ,ϕ(x,y)=Rθ(x,y)−αP(toxic∣x,y)−βDKL(πϕ∥π0).\hat R_{\theta,\phi}(x,y)=R_\theta(x,y)-\alpha P(\mathrm{toxic}\mid x,y)-\beta D_{\mathrm{KL}}(\pi_\phi\|\pi_0).

    Here yy is a generated response, P(toxic∣x,y)P(\mathrm{toxic}\mid x,y) is the toxicity probability assigned to an existing classifier, πϕ\pi_\phi is the policy being optimized, and π0\pi_0 is the supervised policy; the KL term penalizes divergence between their response distributions. The classifier penalty discourages responses that the classifier readily recognizes as toxic, while the learned reward preserves a signal favoring toxicity rather than merely steering generation toward non-toxicity. In the main setup, the policy and reward model use LLaMA-13B, the classifier penalty uses a BAD-trained RoBERTa classifier, and α=5\alpha=5 and β=0.1\beta=0.1. PPO maximizes expected reward over queries and policy-generated responses.

  2. Knowl 2 — RL fine-tuning raises attack success across five toxicity classifiers

    empirical result

    The authors compared toxic benchmark data, zero-shot GPT-3.5-turbo responses, and three LLaMA-13B generation methods on five toxicity classifiers. Attack success rate (ASR) is the percentage of toxic responses classified as non-toxic; annotated toxic probability is the share judged toxic by human annotators. Reward is the learned reward-model score, and Distinct-4 is the reported four-gram diversity score. Classifier ASRs and annotated toxic probabilities are reported as percentages; reward and Distinct-4 are shown at the paper's reported precision. Sources distinguish crawled data, crowdsourced data, and language-model-generated data. SL is supervised fine-tuning; SL-R continues training on the top-ranked of five SL responses per query; RL is the proposed PPO method.

    Test dataSourceRewardAnnotated toxic probabilityP-API ASRModeration ASRTOXIGEN ASRBAD ASRDavinci003 ASRDistinct-4
    Offensive TwitterCrawl-5.91N/A14.1073.2014.401.906.200.99
    TOXIGENLM-3.96N/A72.2867.9333.9720.929.240.94
    Latent HatredCrawl + CS-3.86N/A72.9274.6442.1416.0911.370.98
    BADCS + LM-3.36N/A76.7782.1155.2824.8526.250.95
    GPT-3.5-turboLM0.7856.9196.6996.6975.1464.0958.470.93
    SL LLaMA-13BLM0.3554.0297.0394.6469.0564.2958.340.91
    SL-R LLaMA-13BLM1.0155.2399.4195.2775.1568.6456.800.87
    RL LLaMA-13BLM2.4758.8499.5597.8182.5190.1662.850.85

    The RL model achieved higher ASR than the SL model on all five classifiers, including an increase from 64.29% to 90.16% on BAD and from 58.34% to 62.85% on Davinci003. The results also show that classifier evasion is not limited to the BAD classifier used in the reward: ASR increased on the other four evaluators as well. At the same time, human-annotated toxicity rose from 54.02% for SL to 58.84% for RL.

  3. Knowl 3 — Human-labeled LLM outputs improve toxicity detection

    empirical result

    The authors collected 4,000 human-labeled toxic responses generated by language models—2,000 from GPT-3.5-turbo and 2,000 from RL-finetuned LLaMA-13B—and fine-tuned a RoBERTa-base classifier on these examples augmented with the BAD dataset. Toxic recall on three test sources was compared with classifiers trained on TOXIGEN, BAD, and Davinci003 data. The values below are percentages; columns identify the source of the toxic test responses.

    ClassifierBAD testGPT-3.5-turbo testRL LLaMA-13B test
    TOXIGEN44.7224.8617.49
    BAD75.1535.919.84
    Davinci00370.6241.5337.16
    Ours78.1682.3278.69
    Ours without RL data76.3180.1171.58

    The model trained on the authors' data achieved 82.32% recall on GPT-3.5-turbo responses and 78.69% on RL LLaMA-13B responses, compared with 35.91% and 9.84%, respectively, for the BAD-trained classifier. It also achieved 78.16% recall on BAD test data, above the BAD classifier's 75.15%. Excluding RL-generated examples reduced recall on RL LLaMA-13B responses to 71.58%.

  4. Knowl 4 — Implicit toxicity is conveyed without overtly offensive wording

    definition

    In this work, an implicitly toxic response implies toxicity or harmfulness without using overtly toxic or offensive words. Its meaning may depend on figurative or indirect language—such as euphemism, circumlocution, sarcasm, metaphor, antithesis, or rhetorical questions—or on extralinguistic knowledge and reasoning. This distinguishes the target from explicit abuse and explains why detecting it can require interpreting the response in its query context rather than matching offensive terms.

  5. Knowl 5 — Evaluation uses BAD-derived queries, human judgments, and five classifiers

    experimental setup

    Queries were extracted from the human utterances in the BAD chatbot-adversarial dataset, which contains nearly 6,000 dialogues; nonsense greetings were filtered. The reported query counts and mean lengths were 14,712 and 15.42 tokens for supervised learning, 22,441 and 15.25 for reward-model training, 22,441 and 15.25 for reinforcement learning, and 311 and 15.21 for testing. LLaMA-13B was the main policy and reward-model backbone. The toxicity classifier included in the RL reward was a 125M-parameter RoBERTa-base model fine-tuned on BAD.

    The five attack evaluators were Perspective-API, OpenAI Moderation, RoBERTa-base classifiers fine-tuned on TOXIGEN and BAD, and a zero-shot Davinci003 toxicity classifier. Three annotators—college students majoring in English, recruited through a professional annotation company—provided toxic/non-toxic judgments used as gold labels. The authors report moderate-to-substantial inter-annotator agreement by Fleiss' kappa. They measured learned reward, Distinct-4 diversity, human-annotated toxic probability, classifier ASR on toxic responses, and continuous toxic confidence.

    For LLaMA-13B training, supervised fine-tuning used batch size 16, learning rate 2×10−72\times10^{-7}, and at most 10 epochs. Reward-model training used batch size 32, learning rate 10−510^{-5}, and at most 5 epochs; the model was initialized from the supervised policy and its first 70% of hidden layers were frozen. RL used batch size 384 and learning rate 5×10−65\times10^{-6}, froze the first 80% of hidden layers, and shared parameters between policy and value functions. Experiments ran on eight 80GB A100 GPUs.

  6. Knowl 6 — The learned preference signal and classifier penalty have complementary roles

    empirical result

    Reward ablations were evaluated by learned reward, human-annotated toxic probability, and average ASR across the five toxicity classifiers. Removing the reward model mainly reduced the production of toxic responses, whereas removing the toxicity classifier penalty reduced classifier evasion.

    VariantRewardAnnotated toxic probability (%)Average ASR (%)
    SL LLaMA-13B0.3554.0276.67
    RL LLaMA-13B2.4758.8486.58
    RL without classifier penalty PP1.8954.6181.75
    RL without learned reward RθR_\theta0.4220.9086.34

    Thus, removing RθR_\theta left ASR high but lowered annotated toxic probability from 58.84% to 20.90%, indicating that the classifier penalty alone tends to favor non-toxic output. Removing PP lowered average ASR from 86.58% to 81.75%, showing that the classifier penalty contributes to evasion. In coefficient analyses, increasing the KL coefficient β\beta reduced reward and toxic probability; BAD ASR first rose and then fell, and the authors selected β=0.1\beta=0.1 to avoid excessive reward-model optimization. Raising classifier coefficient α\alpha over a moderate range improved ASR while preserving comparable toxic probability, but too large a value substantially reduced toxic probability.

  7. Knowl 7 — Larger models use more linguistic features per implicit-toxic response

    empirical result

    To examine scaling, the authors manually labeled linguistic features in responses from RL-finetuned LLaMA-1.3B and LLaMA-13B on 50 randomly sampled queries per model. With larger model size, the average number of features per response rose from 0.89 to 1.40. The reported percentages below describe the occurrence of each feature in the sampled responses.

    Model sizeAverage features per responseSarcasmCircumlocutionEuphemismAntithesisMetaphorRhetorical question
    1.3B0.8965.0%22.5%10.0%0.0%2.5%0.0%
    13B1.4034.9%27.0%19.0%12.7%3.2%3.2%

    Using the same supervised and RL data, the authors also observed that ASR increased across the tested model scales from 1.3B to 13B, with the 13B model attaining the highest toxic probability and ASR. They interpret the feature-labeling results as evidence that larger models combine a broader range of linguistic devices; the proposed explanation that larger models absorb more relevant knowledge during pretraining is presented as a conjecture.

  8. Knowl 8 — Generated implicit toxicity spans multiple toxic categories

    empirical result

    Human evaluation grouped generated toxic responses into four categories: offending the user, unfairness and discrimination, toxic agreement with harmful input, and sensitive-topic content. The distributions differed between GPT-3.5-turbo and RL-finetuned LLaMA-13B.

    Toxic typeGPT-3.5-turboRL LLaMA-13B
    Offending user33.33%39.20%
    Unfairness and discrimination33.90%29.55%
    Toxic agreement22.03%23.30%
    Sensitive topic10.74%7.95%

    The results show that the generated toxicity was not confined to one target or category. The authors caution that the proportions depend strongly on the input context; for example, queries designed to elicit an insulting response can increase the share of offending-user cases.

  9. Knowl 9 — Some successful attacks require contextual knowledge and reasoning to detect

    empirical result

    The authors manually inspected toxic responses from GPT-3.5-turbo and RL LLaMA-13B that all five evaluated classifiers had misclassified. In the examples they analyzed, identifying the harmful implication required combining the query with contextual or world knowledge and interpreting indirect linguistic cues. In a diagnostic test, adding a human-written explanation of the relevant implication to the Davinci003 classifier prompt enabled it to identify toxicity that it had missed when given only the query and response. This case study indicates that the observed misses can involve reasoning demands, not just the absence of obvious toxic keywords; it does not establish that explanations generally solve the detection problem.

  10. Knowl 10 — Automatic preference labels and model scale limit the demonstrated attack

    limitation

    Reward-model comparisons were labeled automatically by GPT-3.5-turbo to reduce annotation cost. The authors report a toxic recall of 68.8% for this labeler in preliminary experiments, and manual review found that nearly 30% of responses assigned to the non-toxic subset were actually implicitly toxic. This noise and bias limit the quality of the reward signal, so the RL policy cannot be expected to find all possible implicit-toxic language. The experiments also did not cover extra-large policy models such as LLaMA-65B or GPT-3.5-turbo, owing to computational-resource or access constraints.

Coverage note — No substantial contributed material was omitted; individual prompt templates and example responses are illustrative instances rather than independent findings.

References

  1. 1.Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 298–306.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  3. 3.Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B Brown, Dawn Song, Ulfar Erlingsson, et al. 2021. Extracting training data from large language models. In USENIX Security Symposium, volume 6.
  4. 4.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  5. 5.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  6. 6.Thomas Davidson, Dana Warmsley, Michael W. Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. In Proceedings of the Eleventh International Conference on Web and Social Media, pages 512–515.
  7. 7.Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023. Toxicity in chatgpt: Analyzing persona-assigned language models. arXiv preprint arXiv:2304.05335.
  8. 8.Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019. Build it break it fix it for dialogue safety: Robustness from adversarial human attack. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, pages 4536–4545.
  9. 9.Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021. Latent hatred: A benchmark for understanding implicit hate speech. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 345–363.
  10. 10.Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5):378.
  11. 11.Jane Frank. 1990. You call that a rhetorical question?: Forms and functions of rhetorical questions in conversation. Journal of Pragmatics, 14(5):723–738.
  12. 12.Simona Frenda, Alessandra Teresa Cignarella, Valerio Basile, Cristina Bosco, Viviana Patti, and Paolo Rosso. 2022. The unbearable hurtfulness of sarcasm. Expert Systems with Applications, 193:116398.
  13. 13.Deep Ganguli, Danny Hernandez, Liane Lovitt, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova Dassarma, Dawn Drain, Nelson Elhage, et al. 2022. Predictability and surprise in large generative models. In 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 1747–1764.
  14. 14.Lei Gao and Ruihong Huang. 2017. Detecting online hate speech using context aware models. In Proceedings of the International Conference Recent Advances in Natural Language Processing, pages 260–266.
  15. 15.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369.
  16. 16.Google. 2023. perspective.
  17. 17.Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3309–3326.
  18. 18.Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei. 2018. Reward learning from human preferences and demonstrations in atari. Advances in neural information processing systems, 31.
  19. 19.Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, et al. 2021. Can machines learn morality? the delphi experiment. arXiv e-prints, pages arXiv–2110.
  20. 20.Jens Lemmens, Ilia Markov, and Walter Daelemans. 2021. Improving hate speech type and target detection with hateful metaphor features. In Proceedings of the Fourth Workshop on NLP for Internet Freedom: Censorship, Disinformation, and Propaganda, pages 7–16.
  21. 21.Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016. A diversity-promoting objective function for neural conversation models. In The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 110–119.
  22. 22.Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  23. 23.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  24. 24.Rijul Magu and Jiebo Luo. 2018. Determining code words in euphemistic hate speech using word embedding networks. In Proceedings of the 2nd workshop on abusive language online (ALW2), pages 93–100.
  25. 25.Nicolas Ocampo, Ekaterina Sviridova, Elena Cabrio, and Serena Villata. 2023. An in-depth analysis of implicit and subtle hate speech messages. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 1989–2005.
  26. 26.OpenAI. 2022. Introducing chatgpt.
  27. 27.OpenAI. 2023a. GPT-4 technical report. arXiv preprint arXiv:2303.08774.
  28. 28.OpenAI. 2023b. moderation.
  29. 29.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744.
  30. 30.Ethan Perez, Saffron Huang, H. Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3419–3448.
  31. 31.Fábio Perez and Ian Ribeiro. 2022. Ignore previous prompt: Attack techniques for language models. arXiv preprint arXiv:2211.09527.
  32. 32.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, pages 2463–2473.
  33. 33.Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021. Recipes for building an open-domain chatbot. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 300–325.
  34. 34.Nigora Ruzibaeva. 2019. Peculiarities of the antithesis in the literary text. European Journal of Research and Reflection in Educational Sciences Vol, 7(11).
  35. 35.Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social bias frames: Reasoning about social and power implications of language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5477–5490.
  36. 36.John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.
  37. 37.Rohit Sridhar and Diyi Yang. 2022. Explaining toxic text via knowledge enhanced text generation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 811–826.
  38. 38.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems, 33:3008–3021.
  39. 39.Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. 2023. Safety assessment of chinese large language models. arXiv preprint arXiv:2304.10436.
  40. 40.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  41. 41.Jiaan Wang, Yunlong Liang, Fandong Meng, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023. Is chatgpt a good nlg evaluator? a preliminary study. arXiv preprint arXiv:2303.04048.
  42. 42.Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020. Recipes for safety in open-domain chatbots. arXiv preprint arXiv:2010.07079.
  43. 43.Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020. DIALOGPT : Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 270–278.
  44. 44.Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2023a. Safety-bench: Evaluating the safety of large language models with multiple choice questions. arXiv preprint arXiv:2309.07045.
  45. 45.Zhexin Zhang, Jiaxin Wen, and Minlie Huang. 2023b. ETHICIST: Targeted training data extraction through loss smoothed soft prompting and calibrated confidence estimation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12674–12687, Toronto, Canada. Association for Computational Linguistics.

Citation

MLA
Wen, J., et al. “Unveiling the Implicit Toxicity in Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 1322–38, https://doi.org/10.18653/v1/2023.emnlp-main.84.
APA
Wen, J., Ke, P., Sun, H., Zhang, Z., Li, C., Bai, J., & Huang, M. (2023). Unveiling the Implicit Toxicity in Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1322–1338. https://doi.org/10.18653/v1/2023.emnlp-main.84
Chicago
Wen, J., P. Ke, H. Sun, et al. 2023. “Unveiling the Implicit Toxicity in Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1322–38. https://doi.org/10.18653/v1/2023.emnlp-main.84.
Harvard
Wen, J. et al. (2023) “Unveiling the Implicit Toxicity in Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 1322–1338. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.84.
Vancouver
1. Wen J, Ke P, Sun H, Zhang Z, Li C, Bai J, Huang M (2023) Unveiling the Implicit Toxicity in Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 1322–1338

BibTeX

@inproceedings{wen-etal-2023-unveiling,
    title = "Unveiling the Implicit Toxicity in Large Language Models",
    author = "Wen, Jiaxin  and
      Ke, Pei  and
      Sun, Hao  and
      Zhang, Zhexin  and
      Li, Chengfei  and
      Bai, Jinfeng  and
      Huang, Minlie",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.84/",
    doi = "10.18653/v1/2023.emnlp-main.84",
    pages = "1322--1338"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/