Self-Detoxifying Language Models via Toxification Reversal

Chak Tou LeongYi ChengJiashuo WangJian WangWenjie Li

article2023EMNLP67 citations

Proposes an inference-time self-detoxification method that steers pretrained language models away from harmful text by identifying and reversing toxification representations within attention layers without requiring model fine-tuning or external classifiers.

Listen

Large language models frequently absorb offensive, disrespectful, or harmful text from their broad web-training data, creating serious safety and compliance risks when deployed in commercial applications. Existing solutions to mitigate this behavior typically rely on fine-tuning entire models on curated datasets or training external classifier modules to filter outputs at runtime. However, fine-tuning requires massive computational resources and risks degrading general performance, while external classifiers significantly increase memory overhead, introduce latency, and often harm output fluency.

The article demonstrates a lightweight "self-detoxification" method that reduces toxic text generation without requiring parameter fine-tuning, auxiliary models, or separate classifier components. The approach identifies how harmful prompts steer internal attention mechanisms and then reverses that direction during normal text generation to suppress harmful output.

The authors evaluated the framework on the RealToxicityPrompts benchmark using a standard base model (GPT-2 Large), testing performance across thousands of non-toxic and toxic prompts. The method operates during inference using two forward passes per generated token: the first pass uses paired positive and negative steering prefixes to discover the "toxification direction" across multi-head attention layers, and the second pass applies an adaptive, scaled reversal to the original prompt's internal representation vectors. The article measured toxic risk using an offline classifier and assessed fluency and relevance using perplexity metrics and human evaluation.

The evaluation revealed several key findings. First, on non-toxic prompts, the method cut the probability of generating toxic text by more than half (dropping from 38.2% in the base model to 17.5%) while preserving high fluency, outperforming alternative prompt-based and fine-tuning baselines. Second, in direct human evaluations, the proposed method was judged less toxic than fine-tuned and decoding-based alternatives by a winning margin of more than two-to-one, while matching them in coherence and fluency. Third, internal layer analysis showed that detoxification occurs predominantly in the middle-to-upper layers (above layer 16), whereas lower layers contribute minimally. Finally, runtime benchmarks demonstrated that this internal steering achieves lower latency (828 ms per sample) and requires no additional model parameters compared to state-of-the-art decoding systems that demand up to three times more parameters.

These findings indicate that organizations can achieve effective safety guardrails at substantially lower infrastructure and operating costs by steering a model's internal representation rather than maintaining external classifier pipelines or executing expensive retraining cycles. Because the method does not modify final token probabilities externally, it avoids the typical trade-off where safety interventions compromise text naturalness. However, decision-makers should note key technical boundaries: the technique requires internal white-box access to model weights—making it unsuitable for closed commercial APIs—and its effectiveness relies on the model's pre-existing internal associations with the steering prompts.

Organizations operating open-weight language models should consider piloting this internal representation reversal as an efficient, inference-time safety layer. Future work should focus on testing this technique on larger and modern instruction-tuned architectures, exploring automated optimization of steering prompts, and investigating whether skipping modifications in the bottom layers can further accelerate real-time serving pipelines.

Cover for Self-Detoxifying Language Models via Toxification Reversal

Abstract

Language model detoxification aims to minimize the risk of generating offensive or harmful content in pretrained language models (PLMs) for safer deployment. Existing methods can be roughly categorized as finetuning-based and decoding-based. However, the former is often resource-intensive, while the latter relies on additional components and potentially compromises the generation fluency. In this paper, we propose a more lightweight approach that enables the PLM itself to achieve “self-detoxification”. Our method is built upon the observation that prepending a negative steering prompt can effectively induce PLMs to generate toxic content. At the same time, we are inspired by the recent research in the interpretability field, which formulates the evolving contextualized representations within the PLM as an information stream facilitated by the attention layers. Drawing on this idea, we devise a method to identify the toxification direction from the normal generation process to the one prompted with the negative prefix, and then steer the generation to the reversed direction by manipulating the information movement within the attention layers. Experimental results show that our approach, without any fine-tuning or extra components, can achieve comparable performance with state-of-the-art methods.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 3 Method
  • 3.1 Toxification Direction Discovery
  • 3.2 Adaptive Toxification Reversal
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Automatic Evaluation
  • 4.3 Human Evaluation
  • 5 Analysis
  • 5.1 Layer-wise Ablation Study
  • 5.2 Analysis on Head-wise Scaling Factors
  • 5.3 Analysis on Detoxification Dynamics
  • 6 Related Works
  • 7 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Baselines
  • B Offline Toxicity Scorer
  • C Effect of Different Scaling Strategies
  • D Discussion on Prefix
  • E Cases of Detoxification Dynamic
  • F Additional Comparison with Text Detoxification
  • G Discussion on Computational Cost

Knowls

  1. Knowl 1 — Contrastive prompts identify each attention head’s toxification direction

    model/method

    At each token-generation step, the method compares the last-token value vectors produced by a negative-prefixed and a positive-prefixed version of the same context. The negative prefix used in the experiments was The following text is abusive, harmful, negative, obscene, racist, rude and toxic:, and the positive prefix was The following text is kind, polite, positive, respectful and supportive:. For transformer layer ℓ\ell and attention head hh, let vN−−,(ℓ,h)v_{N^-}^{-,(\ell,h)} and vN++,(ℓ,h)v_{N^+}^{+,(\ell,h)} be the contextualized value vectors at the final positions of the respective prefixed inputs, whose lengths are N−N^- and N+N^+. The head-specific toxification direction is

    Δv(ℓ,h)=vN−−,(ℓ,h)−vN++,(ℓ,h).\Delta v^{(\ell,h)} = v_{N^-}^{-,(\ell,h)} - v_{N^+}^{+,(\ell,h)}.

    In a second forward pass on the original, unprefixed context, let vN(ℓ,h)v_N^{(\ell,h)} be the current value vector at its last position. The basic reversal subtracts the discovered direction, vNnew,(ℓ,h)=vN(ℓ,h)−Δv(ℓ,h)v_N^{\mathrm{new},(\ell,h)}=v_N^{(\ell,h)}-\Delta v^{(\ell,h)}. Only the last-position representation is changed, leaving earlier prompt representations intact to avoid altering the context. The paired positive prompt provides a contrasting reference intended to isolate the negative prompt’s toxification effect while reducing interference with contextual semantics.

  2. Knowl 2 — Adaptive reversal scales the intervention and preserves the layer’s total value-vector norm

    equation

    The method strengthens reversal for heads with a larger discovered toxification direction and for current value vectors similar to the negative-prefix representation. For layer ℓ\ell and head hh, define ∥⋅∥2\|\cdot\|_2 as the Euclidean norm and cos⁡(u,v)=u⋅v/(∥u∥2∥v∥2)\cos(u,v)=u\cdot v/(\|u\|_2\|v\|_2). Let vK−−,(ℓ,h)v_{K^-}^{-,(\ell,h)} denote the contextualized value vector at the final token of the negative prefix, and let vN(ℓ,h)v_N^{(\ell,h)} be the current value vector at the last position of the unprefixed context. The two scaling factors and the adaptive update are

    λnorm=1+∥Δv(ℓ,h)∥2,λsim=1+max⁡{0,cos⁡(vN(ℓ,h),vK−−,(ℓ,h))},\lambda_{\mathrm{norm}}=1+\|\Delta v^{(\ell,h)}\|_2, \qquad \lambda_{\mathrm{sim}}=1+\max\{0,\cos(v_N^{(\ell,h)},v_{K^-}^{-,(\ell,h)})\}, v~N(ℓ,h)=vN(ℓ,h)−λnormαλsimβΔv(ℓ,h).\widetilde v_N^{(\ell,h)}=v_N^{(\ell,h)}-\lambda_{\mathrm{norm}}^{\alpha}\lambda_{\mathrm{sim}}^{\beta}\Delta v^{(\ell,h)}.

    Here v~N(ℓ,h)\widetilde v_N^{(\ell,h)} is the updated head vector before renormalization; the experiments used α=0.4\alpha=0.4 and β=0.6\beta=0.6. After updating all HH heads in a layer, concatenate their vectors as v~N(ℓ)=[v~N(ℓ,1);…;v~N(ℓ,H)]\widetilde v_N^{(\ell)}=[\widetilde v_N^{(\ell,1)};\ldots;\widetilde v_N^{(\ell,H)}], and concatenate the original vectors as vN(ℓ)=[vN(ℓ,1);…;vN(ℓ,H)]v_N^{(\ell)}=[v_N^{(\ell,1)};\ldots;v_N^{(\ell,H)}]. The method rescales the updated concatenation to the original layer-wise norm:

    vNnew,(ℓ)=v~N(ℓ)∥vN(ℓ)∥2∥v~N(ℓ)∥2.v_N^{\mathrm{new},(\ell)}=\widetilde v_N^{(\ell)}\frac{\|v_N^{(\ell)}\|_2}{\|\widetilde v_N^{(\ell)}\|_2}.

    This renormalization is intended to limit disruption to the model’s ordinary representations and preserve generation fluency.

  3. Knowl 3 — Inference requires paired discovery and intervention passes for each generated token

    algorithm

    The procedure operates on a causal language model whose attention-head value vectors can be accessed and modified during inference; it requires no fine-tuning or auxiliary toxicity model. For each next-token prediction, run the model on a batch containing the current context with the negative prefix and with the positive prefix. At every layer and head, retain the difference between the value vectors at the two inputs’ final positions as that head’s toxification direction. Then run a second forward pass on the original context. At each layer, apply the adaptive update and layer-wise renormalization defined in this method’s equations to the value vectors at the context’s final position only. Use the resulting model distribution to sample the next token and append it to the context. Repeat both passes for each subsequent token. The experiments used the two prefixes stated in the contrastive-direction knowl, with α=0.4\alpha=0.4 and β=0.6\beta=0.6.

  4. Knowl 4 — Evaluation uses RealToxicityPrompts with separate initially toxic and non-toxic contexts

    experimental setup

    The evaluation used RealToxicityPrompts, which contains 100,000 web-text paragraphs with toxicity annotations for their prompt portions. The researchers randomly sampled 10,000 prompts, retained the 9,907 with annotations, and evaluated 7,785 prompts scoring below 0.5 as the non-toxic-prompt setting and 2,122 scoring above 0.5 as the toxic-prompt setting. For each prompt, models generated 25 continuations using GPT-2 Large and nucleus sampling with p=0.9p=0.9; each continuation contained between 5 and 20 tokens. Expected Maximum Toxicity is the average, across prompts, of the highest toxicity score among the 25 continuations. Toxicity Probability is the fraction of prompts for which at least one of the 25 continuations scores at least 0.5. Perplexity, used as a fluency measure, was evaluated with GPT-2 XL. Generated-text toxicity was estimated by a DeBERTa-v3-large scorer trained on 90,000 RealToxicityPrompts samples to fit Perspective API scores; on a held-out 10,000-sample set, it achieved 94.87% accuracy and 98.54% AUROC.

  5. Knowl 5 — Automatic evaluation shows lower toxicity than the base model and prompt baseline, with a fluency trade-off against some alternatives

    empirical result

    In the RealToxicityPrompts evaluation, the proposed method reduced toxicity relative to GPT-2 Large in both prompt settings. For non-toxic prompts, its Expected Maximum Toxicity was 0.329, Toxicity Probability was 17.5%, and perplexity was 13.14, compared with the base model’s 0.457, 38.2%, and 11.29. For toxic prompts, the corresponding values were 0.607, 62.5%, and 13.77, versus the base model’s 0.759, 84.2%, and 11.85. It also improved on the strongest reported Self-Debiasing prompt setting (λ=100\lambda=100): non-toxic prompts yielded 0.329, 17.5%, and 13.14 for the proposed method versus 0.355, 20.3%, and 21.09; toxic prompts yielded 0.607, 62.5%, and 13.77 versus 0.623, 65.5%, and 23.32. Results were mixed against other method families: on non-toxic prompts the proposed method’s 0.329 Expected Maximum Toxicity and 17.5% Toxicity Probability were close to DAPT’s 0.331 and 18.9%, while perplexity was lower (13.14 versus 19.72); on toxic prompts DAPT had lower toxicity scores (0.558 and 57.0% versus 0.607 and 62.5%), but higher perplexity (22.47 versus 13.77). DExperts and GeDi reported lower automatic toxicity metrics than the proposed method, but used larger models; the paper cautions that their toxicity scores may be inflated by misclassification from the automatic evaluator. The page 5 evaluation table also compares these methods across both prompt settings and all three metrics. In an additional comparison, generating with the proposed method produced 0.329 Expected Maximum Toxicity, 17.5% Toxicity Probability, and 13.14 perplexity, while applying GPT-2+BART-detox-base to generated text produced 0.428, 34.1%, and 32.87.

  6. Knowl 6 — Human judgments favor the method’s toxicity reduction while usually finding fluency and coherence comparable

    empirical result

    Three graduate-student evaluators compared continuations for 150 sampled prompts: 50 comparisons each against DAPT, DExperts, and Self-Debiasing. For each comparison they judged which continuation was less toxic, more fluent, and more coherent with the prompt, or selected a tie. In win/tie/loss percentages for the proposed method, the less-toxic judgments were 26.8/63.4/9.8 against DAPT, 18.8/71.8/9.4 against DExperts, and 25.6/64.3/10.1 against Self-Debiasing. Thus, for all three comparisons, the proposed method’s less-toxic win rate exceeded its loss rate. Fluency and coherence were mostly judged ties against DAPT and DExperts; against Self-Debiasing, the proposed method won 38.8% of fluency judgments and 52.7% of coherence judgments, with tie/loss rates of 46.5/14.7 and 28.7/18.6, respectively. The human-evaluation chart on page 5 displays these win/tie/loss proportions. Fleiss’s κ\kappa was 0.244, which the authors characterize as fair inter-annotator agreement.

  7. Knowl 7 — Ablation indicates that middle-upper transformer layers contribute most to toxicity reduction

    empirical result

    The layer ablation study removed the reversal operation from groups of layers in GPT-2 Large, either starting at the bottom of the model, starting at the top, or within a four-layer block in the middle. The reported pattern is uneven across depth: removing reversal from middle-lower layers (below roughly layer 16) caused only a small loss in toxicity reduction, and using those layers alone also produced little reduction. By contrast, reversal in the middle-upper layers made a substantial contribution. The layer-ablation plot on page 6 visualizes Expected Maximum Toxicity as the removed layer ranges vary; the authors interpret its pattern as evidence that the detoxifying effect is concentrated more strongly in middle-upper than middle-lower layers.

  8. Knowl 8 — Head-wise scaling correlates with toxicity reduction most clearly for the norm-based factor

    empirical result

    To examine the adaptive factors, the researchers sampled 1,000 non-toxic prompts and generated 25 continuations per prompt from both the base model and the detoxified model. For each prompt they calculated toxicity reduction as the difference between the two models’ average continuation toxicity, then correlated that quantity across prompts with each head’s average scaling factor during generation. The page 7 Spearman-correlation heatmaps show generally higher correlations for λnorm\lambda_{\mathrm{norm}} in middle-upper layers than in middle-lower layers, consistent with the layer-ablation pattern. High correlations for λsim\lambda_{\mathrm{sim}} are sparser, matching the authors’ observation that changing the similarity-based factor affected toxicity less than changing the norm-based factor. They suggest, without establishing it as a demonstrated mechanism, that especially correlated heads may carry toxicity-relevant style or semantic information.

  9. Knowl 9 — Logit-lens analysis shows negative-token probabilities falling during reversal

    empirical result

    For the context Put your hand in my face and I’m going to, the base model tended to predict a violent or toxic continuation. The researchers tracked next-token probabilities across layers using a logit lens and selected 14 candidate negative verbs: slap, beat, break, fuck, hit, hurt, kick, kill, knock, punch, rape, rip, shoot, and smash. In the base model, the combined probability of these verbs rose to nearly 100% around layer 24 and remained above 20% at the output. With toxification reversal, their combined probability was suppressed to nearly zero around layer 16. The probability of slap reached about 4% at the base-model output and was likewise suppressed around layer 16 with reversal. The layer-by-layer probability plots on page 8 show this contrast; the authors note that suppression begins in the region they also identify as important for detoxification.

  10. Knowl 10 — Inference latency is below DExperts while the method adds no parameters or training

    empirical result

    The method adds no learned parameters and requires no additional training. The parameter comparison reports 774 million parameters for GPT-2 Large and the proposed method, versus 2,322 million for DExperts and 1,129 million for GeDi. On 100 randomly sampled prompts with batch size 1 on an NVIDIA 3090 GPU, reported latency was 943±12943\pm12 ms per sample for DExperts, 828±14828\pm14 ms for the proposed method, and 756±14756\pm14 ms for a variant that omitted reversal in the bottom 16 layers. The authors report only marginal toxicity-performance decay for that reduced-layer variant. The page 16 tables report the parameter and latency comparisons; no GeDi latency is given.

  11. Knowl 11 — The method depends on toxicity knowledge in the model and direct access to its internal representations

    limitation

    Toxification reversal can suppress only harmful concepts or forms that the negative prefix can evoke from the pretrained model. Its effectiveness therefore depends on the pretraining data, model training, and model capacity; harmful content not associated with the prefix in the model’s learned representations may remain, and the authors caution that smaller models may be unsuitable. The intervention also requires modifying internal representations during the forward pass, so it cannot be applied through a model API that withholds access to those representations. The authors additionally acknowledge the risks of under-detoxification, over-detoxification that removes valid content, and introducing new biases.

Coverage note — The extensive individual sampled continuations and token-specific plots beyond the summarized negative-verb dynamics are omitted because they illustrate, rather than add distinct findings to, the reported method and analyses.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer normalization.
  2. 2.Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023. Eliciting latent predictions from transformers with the tuned lens.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  4. 4.Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2022. Analyzing transformers in embedding space.
  5. 5.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  6. 6.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread. Https://transformer-circuits.pub/2021/framework/index.html.
  7. 7.Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I. Liao, Kamile Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, Dawn Drain, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jackson Kernion, Jamie Kerr, Jared Mueller, Joshua Landau, Kamal Ndousse, Karina Nguyen, Liane Lovitt, Michael Sellitto, Nelson Elhage, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robert Lasenby, Robin Larson, Sam Ringer, Sandipan Kundu, Saurav Kadavath, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, Christopher Olah, Jack Clark, Samuel R. Bowman, and Jared Kaplan. 2023. The capacity for moral self-correction in large language models.
  8. 8.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The pile: An 800gb dataset of diverse text for language modeling.
  9. 9.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3356–3369, Online. Association for Computational Linguistics.
  10. 10.Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.
  11. 11.Aaron Gokaslan and Vanya Cohen. 2019. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus.
  12. 12.Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360, Online. Association for Computational Linguistics.
  13. 13.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023. DeBERTav3: Improving deBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing. In The Eleventh International Conference on Learning Representations.
  14. 14.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations.
  15. 15.Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019. Ctrl: A conditional transformer language model for controllable generation.
  16. 16.Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, and Ethan Perez. 2023. Pre-training language models with human preferences.
  17. 17.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2021. GeDi: Generative discriminator guided sequence generation. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 4929–4952, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  18. 18.Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang. 2022. Fine-tuning can distort pretrained features and underperform out-of-distribution. In International Conference on Learning Representations.
  19. 19.Sachin Kumar, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos, and Yulia Tsvetkov. 2023. Language generation models can cause harm: So what can we do about it? an actionable survey.
  20. 20.Jin Myung Kwak, Minseon Kim, and Sung Ju Hwang. 2022. Language detoxification with attribute-discriminative latent space.
  21. 21.Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021. DExperts: Decoding-time controlled text generation with experts and anti-experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6691–6706, Online. Association for Computational Linguistics.
  22. 22.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Comput. Surv., 55(9).
  23. 23.Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022. ParaDetox: Detoxification with parallel data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6804–6818, Dublin, Ireland. Association for Computational Linguistics.
  24. 24.Ximing Lu, Sean Welleck, Jack Hessel, Liwei Jiang, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi. 2022. Quark: Controllable text generation with reinforced unlearning. In Advances in Neural Information Processing Systems, volume 35, pages 27591–27609. Curran Associates, Inc.
  25. 25.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730–27744. Curran Associates, Inc.
  26. 26.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only.
  27. 27.Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2023. Adding instructions during pretraining: Effective way of controlling toxicity in language models. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2636–2651, Dubrovnik, Croatia. Association for Computational Linguistics.
  28. 28.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners.
  29. 29.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(1).
  30. 30.Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. 2023. Toward transparent ai: A survey on interpreting the inner structures of deep neural networks. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pages 464–483.
  31. 31.Timo Schick, Sahana Udupa, and Hinrich Schütze. 2021. Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP. 9:1408–1424.
  32. 32.Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020. Investigating gender bias in language models using causal mediation analysis. Advances in neural information processing systems, 33:12388–12401.
  33. 33.Boxin Wang, Wei Ping, Chaowei Xiao, Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Bo Li, Anima Anandkumar, and Bryan Catanzaro. 2022. Exploring the limits of domain-adaptive training for detoxifying large-scale language models. In Advances in Neural Information Processing Systems.
  34. 34.Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V Le. 2022. Finetuned language models are zero-shot learners. In International Conference on Learning Representations.
  35. 35.Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein. 2021. Detoxifying language models risks marginalizing minority voices. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2390–2397, Online. Association for Computational Linguistics.
  36. 36.Canwen Xu, Zexue He, Zhankui He, and Julian McAuley. 2022. Leashing the inner demons: Self-detoxification for language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11530–11537.
  37. 37.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. 2023. Representation engineering: A top-down approach to ai transparency.

Citation

MLA
Leong, C. T., et al. “Self-Detoxifying Language Models via Toxification Reversal”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 4433–49, https://doi.org/10.18653/v1/2023.emnlp-main.269.
APA
Leong, C. T., Cheng, Y., Wang, J., Wang, J., & Li, W. (2023). Self-Detoxifying Language Models via Toxification Reversal. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4433–4449. https://doi.org/10.18653/v1/2023.emnlp-main.269
Chicago
Leong, C. T., Y. Cheng, J. Wang, J. Wang, and W. Li. 2023. “Self-Detoxifying Language Models via Toxification Reversal”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4433–49. https://doi.org/10.18653/v1/2023.emnlp-main.269.
Harvard
Leong, C.T. et al. (2023) “Self-Detoxifying Language Models via Toxification Reversal”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4433–4449. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.269.
Vancouver
1. Leong CT, Cheng Y, Wang J, Wang J, Li W (2023) Self-Detoxifying Language Models via Toxification Reversal. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4433–4449

BibTeX

@inproceedings{leong-etal-2023-self,
    title = "Self-Detoxifying Language Models via Toxification Reversal",
    author = "Leong, Chak Tou  and
      Cheng, Yi  and
      Wang, Jiashuo  and
      Wang, Jian  and
      Li, Wenjie",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.269/",
    doi = "10.18653/v1/2023.emnlp-main.269",
    pages = "4433--4449"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/