Reconstructive Neuron Pruning for Backdoor Defense
Yige LiXixiang LyuXingjun MaNodens KorenLingjuan LyuBo LiYu-Gang Jiang
Proposes an asymmetric unlearning and filter-recovery framework that exposes and prunes backdoor neurons using only a small set of clean data, effectively purifying backdoored models across diverse attacks without sacrificing clean classification accuracy.
Deep neural networks are increasingly deployed in mission-critical applications, yet they remain highly vulnerable to backdoor attacks. These malicious attacks manipulate training data or procedures to insert hidden triggers that cause models to misclassify specific inputs while appearing to operate normally on clean data. As organizations rely more on outsourced training and pre-trained models from third parties, the security risks posed by covert backdoors continue to grow. Existing defense mechanisms struggle to cleanly remove these vulnerabilities without degrading overall model accuracy or failing against sophisticated feature-level manipulation.
The article develops and evaluates a post-training defense framework called Reconstructive Neuron Pruning (RNP). Its primary objective is to reliably expose, isolate, and prune backdoor-associated neurons from compromised neural networks using only a minimal set of clean defense samples, while preserving standard task accuracy.
To achieve this, the authors designed an asymmetric two-step reconstruction process evaluated across 12 advanced input-space, feature-space, and adaptive backdoor attacks. Using benchmark image datasets (CIFAR-10, ImageNet subsets, and GTSRB) and standard network architectures such as ResNet, the defense operates using only 500 clean samples (around 1% of training data). The method first unlearns the network at the fine-grained neuron level by maximizing error on the clean defense samples, which temporarily deactivates clean features while leaving dormant backdoor neurons intact. It then recovers performance at the coarser filter level by minimizing error on the same clean data, forcing the model to repurpose the remaining backdoor neurons to restore clean functionality. The repurposed backdoor neurons are then precisely identified and pruned via dynamic thresholding.
The experimental findings show that RNP establishes a new state of the art in backdoor defense. On the CIFAR-10 benchmark, RNP reduced the average attack success rate from 96.34% to 5.03% across 12 diverse attacks, maintaining a high clean accuracy of 92.18% (less than a 2% drop from the unattacked baseline). In comparison, leading alternative defenses left average attack success rates between 13.65% and 56.45%. On high-resolution ImageNet data, RNP reduced the attack success rate from 93.89% to 8.87%, significantly outperforming the next-best method's 18.43%. Furthermore, RNP achieved effective neutralization by pruning roughly half as many neurons as competing approaches (for instance, removing only 41 neurons to defeat standard BadNets attacks). In addition, intermediate models generated during the unlearning stage successfully boosted downstream defense capabilities, enabling a 100% detection rate for target backdoor labels and improving trigger recovery.
These results demonstrate that asymmetric unlearning and recovery is highly practical and cost-effective, eliminating the need to retrain compromised models from scratch. Organizations can secure externally acquired machine learning models quickly and with minimal validated clean data, significantly reducing cybersecurity risks and computational costs in automated vision systems.
For practical deployment, organizations should adopt network-wide filter pruning using dynamic thresholding when maximum security against attacks is the primary operational concern. Alternatively, applying pruning selectively to the network's deepest layers offers a balanced option that fully preserves clean accuracy while still removing the vast majority of backdoor vulnerabilities. Operators should avoid excessive unlearning epochs, which risk model collapse, and should monitor validation loss closely.
Confidence in the reported results is high for standard convolutional architectures and moderately poisoned networks. However, decision-makers should note three key operational limitations: RNP's defensive performance degrades significantly on lightweight architectures like EfficientNet (where clean accuracy drops to around 65%), effectiveness decreases when attacks use extremely low poisoning rates (0.05% or less), and RNP does not provide theoretical security guarantees against entirely novel, unforeseen attack designs.
- Paper: Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks, Kang Liu et al. (2018). Introduces the foundational fine-pruning defense against neural network backdoors that RNP builds upon and directly outperforms through its asymmetric neuron unlearning and recovery mechanism.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). Establishes the foundational BadNets threat model and attack mechanism evaluated as primary benchmark threats by the defense framework in this paper.
- Paper: Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning, Xinyun Chen et al. (2017). Provides the essential framework for targeted data-poisoning backdoor attacks in deep neural networks that RNP is designed to neutralize.
- Paper: DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints, Zhendong Zhao et al. (2022). Develops latent representation and feature-space backdoor attacks that represent the sophisticated class of attacks RNP specifically aims to detect and prune.
- Paper: Handcrafted Backdoors in Deep Neural Networks, Sanghyun Hong et al. (2022). Examines how backdoor triggers manipulate specific internal neurons and filters, providing the mechanistic basis for neuron-level isolation and pruning.
- Paper: Learning both Weights and Connections for Efficient Neural Networks, Song Han et al. (2015). Presents foundational neuron and connection pruning techniques in deep neural networks that underlie structural pruning defense mechanisms.
- Paper: Detecting Backdoors in Pre-trained Encoders, Shiwei Feng et al. (2023). Extends backdoor detection from supervised classifiers to pre-trained, self-supervised vision encoders without requiring downstream labeled data or classifier heads.
- Paper: Backdoor Attacks Against Deep Image Compression via Adaptive Frequency Trigger, Yi Yu et al. (2023). Explores how backdoor vulnerabilities and trigger mechanisms propagate beyond high-level classification into low-level image processing models like deep image compression.
