keyword
backdoor removal methods
Backdoor removal methods are machine learning security techniques designed to eliminate hidden malicious triggers and behaviors from compromised neural networks while preserving their standard performance on benign data. When an artificial intelligence model is poisoned during training, it can function normally on standard inputs yet produce attacker-specified outputs whenever a specific trigger pattern is present. Backdoor removal methods sanitize these infected models post-training using strategies such as fine-tuning, machine unlearning, knowledge distillation, trigger inversion, and neuron pruning to detect and deactivate backdoor-related parameters. By repairing the network weights and architecture using a limited set of verified clean data, these defenses restore model safety and integrity without the need to retrain the entire network from scratch.
1 item

