keyword
self-detoxifying language models
Self-detoxifying language models are artificial intelligence text-generation systems capable of autonomously mitigating and preventing the output of toxic, offensive, or harmful language using their own internal mechanisms. Unlike conventional detoxification techniques that rely on resource-intensive model fine-tuning or separate auxiliary filter models during decoding, self-detoxifying models leverage their existing internal representations and activation pathways to redirect generation away from harmful patterns. By internally identifying and reversing toxic tendencies during the generation process, these models reduce unsafe outputs while preserving generation fluency, model capabilities, and computational efficiency.
1 item

