keyword
language model detoxification
Language model detoxification refers to the process of reducing or eliminating toxic, offensive, abusive, and harmful text generated by pretrained language models to ensure safer deployment. Because language models are typically trained on vast, uncurated web corpora, they risk internalizing and reproducing profanity, hate speech, bias, and harassment. Detoxification strategies address this vulnerability through several methods, including filtering training datasets, fine-tuning model parameters using safety-aligned data or reinforcement learning, guiding decoding algorithms during inference, and steering internal model representations. The primary goal of detoxification is to suppress unsafe text generation while preserving the model linguistic fluency, factual accuracy, and general task performance.
1 item

