keyword
neuron unlearning
Neuron unlearning is a machine learning process that selectively removes or suppresses specific memorized information, malicious patterns, or unintended behaviors from a trained neural network by directly altering the weights, activations, or representations of individual neurons. Rather than retraining an entire model from scratch, this technique modifies the internal neural components associated with targeted data—such as backdoor triggers, sensitive user records, or erroneous associations—often by perturbing, masking, or optimizing neuron parameters to degrade performance on unwanted inputs while preserving accuracy on valid data. By operating at the granularity of individual units within network layers, neuron unlearning provides a targeted and computationally efficient mechanism for model editing, privacy compliance, and vulnerability remediation without compromising overall network functionality.
1 item

