Built independently by an author, for readers. Read the story and support ChapterPal

keyword

toxic vectors

Toxic vectors refer to directional representations within the latent activation or parameter spaces of neural language models that correspond to and promote the generation of harmful, abusive, or offensive content. In machine learning and mechanistic interpretability, these vectors capture how undesirable concepts are structurally encoded across network layers, feed-forward submodules, or individual neurons. By locating and isolating these mathematical directions, researchers can analyze the internal mechanisms of safety alignment algorithms, determine whether safety training permanently eliminates or merely bypasses hazardous behaviors, and perform targeted interventions such as activation steering or pruning to suppress toxic outputs.

1 item