keyword
toxification direction
The toxification direction is a vector in the internal representation space of a language model that captures the shift from standard or benign text generation toward the generation of offensive, toxic, or harmful content. It is typically identified by measuring the divergence in contextual representations and attention layer activations between unconstrained generation and generation guided by toxicity-inducing prompts. By isolating this directional trajectory within the model information stream, internal activations can be steered along the opposite vector, enabling the model to reduce harmful outputs and achieve detoxification without requiring extensive fine-tuning or supplementary decoding components.
1 item

