keyword
Neural Cleanse
Neural Cleanse is a security framework designed to detect and mitigate backdoor or Trojan attacks embedded within deep neural networks. In these attacks, a model is trained or modified to perform normally on standard data but produce an attacker-chosen output whenever a specific trigger is present. Neural Cleanse identifies such vulnerabilities by reverse-engineering the minimal input perturbation needed to force all inputs into each target class, operating on the principle that a compromised class requires a significantly smaller trigger than uncompromised classes. After identifying an infected class through outlier detection across these reconstructed triggers, the framework enables defense and remediation through methods such as input filtering, neuron pruning, or unlearning to neutralize the backdoor without severely impacting the model performance on benign tasks.
1 item

