keyword
backdoor defense
Backdoor defense refers to a set of security techniques and countermeasures designed to detect, prevent, or eliminate hidden malicious vulnerabilities within machine learning models and datasets. In a backdoor attack, an adversary manipulates a model during training or development so that it performs accurately on typical inputs but outputs a compromised, targeted prediction whenever a specific trigger pattern is present. Backdoor defenses mitigate these risks throughout the machine learning lifecycle by filtering poisoned training samples, employing robust training mechanisms, auditing model parameters, and repairing infected networks through methods such as neuron pruning, fine-tuning, or trigger unlearning. Additionally, these defenses can operate during inference to identify and sanitize triggered inputs, ensuring that the system remains secure without sacrificing its predictive performance on benign data.
1 item

