Reconstructive Neuron Pruning for Backdoor Defense

Yige LiXixiang LyuXingjun MaNodens KorenLingjuan LyuBo LiYu-Gang Jiang

article2023ICML84 citations

Proposes an asymmetric unlearning and filter-recovery framework that exposes and prunes backdoor neurons using only a small set of clean data, effectively purifying backdoored models across diverse attacks without sacrificing clean classification accuracy.

Listen

Deep neural networks are increasingly deployed in mission-critical applications, yet they remain highly vulnerable to backdoor attacks. These malicious attacks manipulate training data or procedures to insert hidden triggers that cause models to misclassify specific inputs while appearing to operate normally on clean data. As organizations rely more on outsourced training and pre-trained models from third parties, the security risks posed by covert backdoors continue to grow. Existing defense mechanisms struggle to cleanly remove these vulnerabilities without degrading overall model accuracy or failing against sophisticated feature-level manipulation.

The article develops and evaluates a post-training defense framework called Reconstructive Neuron Pruning (RNP). Its primary objective is to reliably expose, isolate, and prune backdoor-associated neurons from compromised neural networks using only a minimal set of clean defense samples, while preserving standard task accuracy.

To achieve this, the authors designed an asymmetric two-step reconstruction process evaluated across 12 advanced input-space, feature-space, and adaptive backdoor attacks. Using benchmark image datasets (CIFAR-10, ImageNet subsets, and GTSRB) and standard network architectures such as ResNet, the defense operates using only 500 clean samples (around 1% of training data). The method first unlearns the network at the fine-grained neuron level by maximizing error on the clean defense samples, which temporarily deactivates clean features while leaving dormant backdoor neurons intact. It then recovers performance at the coarser filter level by minimizing error on the same clean data, forcing the model to repurpose the remaining backdoor neurons to restore clean functionality. The repurposed backdoor neurons are then precisely identified and pruned via dynamic thresholding.

The experimental findings show that RNP establishes a new state of the art in backdoor defense. On the CIFAR-10 benchmark, RNP reduced the average attack success rate from 96.34% to 5.03% across 12 diverse attacks, maintaining a high clean accuracy of 92.18% (less than a 2% drop from the unattacked baseline). In comparison, leading alternative defenses left average attack success rates between 13.65% and 56.45%. On high-resolution ImageNet data, RNP reduced the attack success rate from 93.89% to 8.87%, significantly outperforming the next-best method's 18.43%. Furthermore, RNP achieved effective neutralization by pruning roughly half as many neurons as competing approaches (for instance, removing only 41 neurons to defeat standard BadNets attacks). In addition, intermediate models generated during the unlearning stage successfully boosted downstream defense capabilities, enabling a 100% detection rate for target backdoor labels and improving trigger recovery.

These results demonstrate that asymmetric unlearning and recovery is highly practical and cost-effective, eliminating the need to retrain compromised models from scratch. Organizations can secure externally acquired machine learning models quickly and with minimal validated clean data, significantly reducing cybersecurity risks and computational costs in automated vision systems.

For practical deployment, organizations should adopt network-wide filter pruning using dynamic thresholding when maximum security against attacks is the primary operational concern. Alternatively, applying pruning selectively to the network's deepest layers offers a balanced option that fully preserves clean accuracy while still removing the vast majority of backdoor vulnerabilities. Operators should avoid excessive unlearning epochs, which risk model collapse, and should monitor validation loss closely.

Confidence in the reported results is high for standard convolutional architectures and moderately poisoned networks. However, decision-makers should note three key operational limitations: RNP's defensive performance degrades significantly on lightweight architectures like EfficientNet (where clean accuracy drops to around 65%), effectiveness decreases when attacks use extremely low poisoning rates (0.05% or less), and RNP does not provide theoretical security guarantees against entirely novel, unforeseen attack designs.

Cover for Reconstructive Neuron Pruning for Backdoor Defense

Abstract

Deep neural networks (DNNs) have been found to be vulnerable to backdoor attacks, raising security concerns about their deployment in mission-critical applications. While existing defense methods have demonstrated promising results, it is still not clear how to effectively remove backdoor-associated neurons in backdoored DNNs. In this paper, we propose a novel defense called Reconstructive Neuron Pruning (RNP) to expose and prune backdoor neurons via an unlearning and then recovering process. Specifically, RNP first unlearns the neurons by maximizing the model’s error on a small subset of clean samples and then recovers the neurons by minimizing the model’s error on the same data. In RNP, unlearning is operated at the neuron level while recovering is operated at the filter level, forming an asymmetric reconstructive learning procedure. We show that such an asymmetric process on only a few clean samples can effectively expose and prune the backdoor neurons implanted by a wide range of attacks, achieving a new state-of-the-art defense performance. Moreover, the unlearned model at the intermediate step of our RNP can be directly used to improve other backdoor defense tasks including backdoor removal, trigger recovery, backdoor label detection, and backdoor sample detection. Code is available at https://github.com/bboylyg/RNP.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Backdoor Attack
  • 2.2. Backdoor Defense
  • 3. Proposed Method
  • 3.1. Threat Model
  • 3.2. Reconstructive Neuron Pruning
  • 3.3. An Illustrative Example
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Main Defense Results
  • 4.3. Understanding the Mechanism of RNP
  • 5. Limitation
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Implementation Details of RNP
  • A.1. Datasets and Classifiers
  • A.2. Attack Details
  • A.3. Defense Details
  • B. Additional Experimental Results of RNP
  • B.1. Comparison between RNP and ANP on the GTSRB Dataset
  • B.2. RNP with Different Pruning Strategies
  • B.3. RNP with Different Model Architectures
  • B.4. RNP against All-to-All Attacks
  • B.5. RNP against Different Trigger Sizes
  • B.6. RNP against Different Poisoning Rate
  • C. More Understandings of RNP
  • C.1. Necessity of Neural Unlearning and Filter Recovering
  • C.2. How Does Filter Location Impact RNP?
  • D. Exploration with the Unlearned Model
  • D.1. Improving Backdoor Removal
  • D.2. Improving Trigger Recovery
  • D.3. Improving Backdoor Label Detection
  • D.4. Improving Backdoor Sample Detection

Knowls

  1. Knowl 1 — Reconstructive Neuron Pruning Algorithm

    algorithm

    Reconstructive Neuron Pruning (RNP) cleanses backdoored deep neural networks by first unlearning clean representations at the neuron level and subsequently recovering clean functionality at the coarser filter level, which isolates and suppresses backdoor neurons.

    Input: Backdoored model fθ(⋅)f_\theta(\cdot) with parameters θ\theta, number of classes KK, clean defense subset Dd\mathcal{D}_d, learning rate η\eta, clean accuracy stopping threshold CAmin⁡CA_{\min}, dynamic pruning threshold DT∈[0,1]DT \in [0, 1]
    Output: Purified model fmκ⊙θf_{m^\kappa \odot \theta}, inferred backdoor target label yty_t
    1: Sample mini-batch (Xd,Yd)(X_d, Y_d) from Dd\mathcal{D}_d
    # Phase 1: Neuron-level unlearning (gradient ascent)
    2: repeat
    3: θ^←arg⁡max⁡θL(f(Xd,Yd;θ))\hat{\theta} \leftarrow \arg\max_\theta \mathcal{L}(f(X_d, Y_d; \theta))
    4: until fθ^f_{\hat{\theta}} clean accuracy on Dd≤CAmin⁡\mathcal{D}_d \le CA_{\min}
    5: Invert/extract backdoor label: yt←arg⁡max⁡Kf(Xd;θ^)y_t \leftarrow \arg\max_K f(X_d; \hat{\theta})
    # Phase 2: Filter-level recovering (filter mask optimization)
    6: Initialize filter mask mκ←[1]nm^\kappa \leftarrow [1]^n
    7: repeat
    8: mκ←mκ−η∂L(f(Xd,Yd;mκ⊙θ^))∂mκm^\kappa \leftarrow m^\kappa - \eta \frac{\partial \mathcal{L}(f(X_d, Y_d; m^\kappa \odot \hat{\theta}))}{\partial m^\kappa}
    9: mκ←clip[0,1](mκ)m^\kappa \leftarrow \text{clip}_{[0, 1]}(m^\kappa)
    10: until filter mask training converges
    # Phase 3: Dynamic threshold pruning
    11: mκ←I(mκ>DT)m^\kappa \leftarrow \mathbb{I}(m^\kappa > DT)
    12: return fmκ⊙θ,ytf_{m^\kappa \odot \theta}, y_t

    In standard implementations, the defense subset Dd\mathcal{D}_d comprises 500 clean samples (1%1\% of training data). The unlearning phase terminates when clean accuracy drops to random guess level (CAmin⁡=10%CA_{\min} = 10\% on a 10-class dataset), typically requiring 5 to 15 epochs with a learning rate of 0.010.01. The filter mask recovery phase runs for 20 epochs with learning rate η=0.2\eta = 0.2. The dynamic threshold DTDT is selected in [0.4,0.7][0.4, 0.7] such that pruning eliminates maximum dormant/repurposed weights while bounding clean accuracy degradation to approximately 2%2\%.

  2. Knowl 2 — Asymmetric Unlearning-Recovering Mechanism for Backdoor Neuron Isolation

    model/method

    Reconstructive Neuron Pruning (RNP) exploits an asymmetry in optimization granularity between an unlearning phase and a recovery phase to expose backdoor neurons using only a small set of clean data Dd\mathcal{D}_d.

    1. Neuron-Level Unlearning (Fine-Grained Maximization): Applying gradient ascent to all individual network weights θ\theta maximizes classification loss on clean data Dd\mathcal{D}_d. Because the defense data contains only clean samples, this process rapidly suppresses the activations of clean-feature neurons while leaving dormant backdoor-associated neurons structurally intact.
    2. Filter-Level Recovering (Coarse-Grained Minimization): Relearning clean features using a continuous mask mκ∈[0,1]nm^\kappa \in [0, 1]^n applied over entire convolutional filters restricts the capacity of the model to directly reverse individual weight changes. To satisfy the classification objective on Dd\mathcal{D}_d under this constraint, the network repurposes the dormant backdoor neurons to compensate for the lost clean features.
    3. Pruning Suppression: Filters that are repurposed to compensate for clean features experience a severe downward scaling in their learned mask values (mκ→0m^\kappa \to 0), whereas filters dedicated to clean features retain high mask values (mκ≈1m^\kappa \approx 1). Applying a threshold mask mκ=I(mκ>DT)m^\kappa = \mathbb{I}(m^\kappa > DT) directly to the original model weights θ\theta purges the backdoor neurons without requiring post-pruning fine-tuning.
  3. Knowl 3 — Mathematical Formulation of Neuron Unlearning and Filter-Mask Recovery

    equation

    Let a backdoored classification model f(⋅;θ)f(\cdot; \theta) have parameter space decomposed as θ=θc∪θb\theta = \theta_c \cup \theta_b, where θc\theta_c denotes clean neurons and θb\theta_b denotes backdoor-associated neurons. Given a clean defense dataset Dd={(xd(i),yd(i))}i=1∣Dd∣\mathcal{D}_d = \{(x_d^{(i)}, y_d^{(i)})\}_{i=1}^{|\mathcal{D}_d|} where ∣Dd∣≪∣D∣|\mathcal{D}_d| \ll |\mathcal{D}|, RNP operates through two sequential optimization objectives.

    First, Neuron Unlearning (NU) optimizes the model parameters θ\theta via empirical loss maximization: θ^=arg⁡max⁡θE(xd,yd)∈DdL(f(xd,yd;θ))\hat{\theta} = \arg\max_{\theta} \mathbb{E}_{(x_d, y_d) \in \mathcal{D}_d} \mathcal{L}(f(x_d, y_d; \theta)) where L\mathcal{L} is the standard cross-entropy loss and θ^\hat{\theta} represents the unlearned parameter state.

    Second, Filter Recovering (FR) freezes θ^\hat{\theta} and optimizes a continuous filter-scaling mask mκ∈[0,1]nm^\kappa \in [0, 1]^n applied element-wise across nn convolutional filters: min⁡mκ∈[0,1]nE(xd,yd)∈DdL(f(xd,yd;mκ⊙θ^))\min_{m^\kappa \in [0, 1]^n} \mathbb{E}_{(x_d, y_d) \in \mathcal{D}_d} \mathcal{L}(f(x_d, y_d; m^\kappa \odot \hat{\theta})) where ⊙\odot denotes channel-wise scaling of the filter parameters by the elements of mκm^\kappa.

  4. Knowl 4 — Defense Effectiveness Against 12 Backdoor Attacks on CIFAR-10 and ImageNet-12

    data/table

    RNP was evaluated against 12 backdoor attacks on CIFAR-10 using ResNet-18 and 5 attacks on an ImageNet-12 subset using ResNet-34. All defenses had access to only 500 clean samples (1%1\% defense data). The evaluated baselines were Fine-Pruning (FP), Neural Attention Distillation (NAD), Implicit Hypergradient Adversarial Unlearning (I-BAU), and Adversarial Neuron Pruning (ANP).

    Dataset / Attack No Defense FP NAD I-BAU ANP RNP (Ours)
    ASR CA ASR CA ASR CA ASR CA ASR CA ASR CA
    CIFAR-10
    BadNets 100.00 93.40 15.11 88.15 1.20 90.64 15.50 91.18 0.53 91.61 0.20 92.22
    Trojan 99.90 93.15 56.51 85.43 5.68 88.72 12.78 90.46 1.00 92.37 2.23 92.56
    Blend 100.00 93.10 68.22 85.21 10.92 89.77 1.62 90.16 0.50 92.31 0.33 92.62
    CL 100.00 94.84 24.38 88.75 13.17 89.76 23.12 88.82 15.20 92.86 8.87 93.12
    SIG 90.86 94.59 19.16 87.88 0.64 89.40 29.32 89.67 1.19 92.97 0.43 94.62
    Dynamic 99.97 91.36 41.73 83.28 13.60 88.64 18.74 86.87 9.20 89.66 15.24 90.18
    WaNet 99.10 93.67 68.92 86.35 17.46 82.41 23.18 87.38 13.14 92.64 10.98 92.83
    FC 100.00 94.67 98.45 87.62 36.07 88.02 17.93 86.75 74.75 81.97 1.80 90.93
    DFST 100.00 94.52 88.78 85.32 12.70 88.72 25.58 87.44 10.80 90.66 4.61 92.78
    AWP 94.39 94.30 23.17 86.04 1.71 89.13 8.71 89.62 0.67 92.64 1.04 94.24
    LIRA 100.00 92.71 87.78 83.12 32.12 86.73 51.33 82.56 20.25 87.78 13.51 92.26
    A-Blend 71.86 92.16 85.22 81.82 18.51 85.23 33.38 85.91 23.71 90.16 1.09 90.38
    Average 96.34 93.54 56.45 85.75 13.65 88.10 21.77 88.07 14.25 90.64 5.03 92.18
    ImageNet-12
    BadNets 100.00 88.53 91.70 83.23 9.12 83.26 15.38 85.15 10.25 85.21 5.80 85.83
    Trojan 100.00 89.79 93.69 81.40 12.31 82.52 19.61 84.11 7.48 87.41 0.59 89.30
    Blend 99.90 89.44 92.14 82.13 28.76 82.93 9.34 82.27 6.21 86.40 5.54 86.89
    SIG 73.78 88.18 87.82 81.27 21.15 83.31 29.23 81.57 25.53 52.52 15.20 84.15
    FC 95.77 88.95 90.52 79.36 31.43 81.56 38.51 79.33 42.69 53.01 17.23 83.36
    Average 93.89 88.98 91.17 81.48 20.55 82.72 22.41 82.49 18.43 72.91 8.87 85.91

    On CIFAR-10, RNP reduces average Attack Success Rate (ASR) from 96.34%96.34\% to 5.03%5.03\% while maintaining 92.18%92.18\% Clean Accuracy (CA), compared to ANP (14.25%14.25\% ASR, 90.64%90.64\% CA) and I-BAU (21.77%21.77\% ASR, 88.07%88.07\% CA). On ImageNet-12, RNP reduces average ASR from 93.89%93.89\% to 8.87%8.87\% with 85.91%85.91\% CA, outperforming ANP which drops CA to 72.91%72.91\%.

  5. Knowl 5 — Neuron Pruning Counts Across Backdoor Attack Types

    data/table

    Comparison of the total number of pruned neurons across all residual blocks of ResNet-18 on CIFAR-10 demonstrates that RNP requires fewer neuron removals than Adversarial Neuron Pruning (ANP) to eliminate backdoors.

    Defense Metric BadNets Trojan Blend CL SIG Dynamic WaNet FC DFST AWP LIRA A-Blend
    ANP Pruned Neurons ↓\downarrow 94 96 42 135 88 69 126 199 165 56 158 96
    ASR (%) 0.53 1.00 0.50 15.20 1.19 9.20 13.14 74.75 10.80 0.67 20.25 23.71
    CA (%) 91.61 92.37 92.31 92.86 92.97 89.66 92.64 81.97 90.66 92.64 87.78 90.16
    RNP Pruned Neurons ↓\downarrow 41 48 28 103 73 59 92 155 83 40 112 78
    ASR (%) 0.20 2.23 0.33 8.87 0.43 15.24 10.98 1.80 4.61 1.04 13.51 1.09
    CA (%) 92.22 92.56 92.62 93.12 94.62 90.18 92.83 90.93 92.78 94.24 92.96 90.38

    For simple input-space attacks (e.g., BadNets, Trojan, Blend), RNP prunes approximately half the number of neurons needed by ANP (41 vs. 94 on BadNets; 48 vs. 96 on Trojan; 28 vs. 42 on Blend). More complex feature-space and adaptive attacks (FC, WaNet, LIRA) distribute trigger associations over broader neuron populations, requiring RNP to prune 78 to 155 neurons.

  6. Knowl 6 — Ablation of Granularity Combinations in Unlearning and Recovering

    empirical result

    Empirical comparison of four combinations of unlearning and recovering on CIFAR-10 with ResNet-18 under BadNets attack demonstrates the necessity of asymmetric granularity:

    1. Neuron Unlearning + Neuron Recovering (NU-NR): Symmetric fine-grained unlearning and recovery restores clean features perfectly but also completely restores backdoor features, failing to reduce Attack Success Rate (ASR remains near 100%100\%).
    2. Filter Unlearning + Filter Recovering (FU-FR): Symmetric coarse-grained unlearning and recovery fails to suppress backdoor activations across both shallow and deep residual blocks.
    3. Filter Unlearning + Neuron Recovering (FU-NR): Coarse unlearning followed by fine-grained recovery mitigates backdoors in deep layers (Block 4) but leaves backdoor activations intact in shallow layers (Block 2).
    4. Neuron Unlearning + Filter Recovering (NU-FR / RNP): Fine-grained neuron unlearning followed by coarse-grained filter recovering suppresses backdoor activations across both shallow and deep layers while restoring clean accuracy, achieving ASR of 0.20%0.20\% and Clean Accuracy of 92.22%92.22\%.
  7. Knowl 7 — Backdoor Label Detection and Trigger Recovery via the Unlearned Model

    model/method

    The intermediate model θ^\hat{\theta} produced by the Neuron Unlearning (NU) step can be directly repurposed to enhance other backdoor defense tasks:

    1. Backdoor Label Detection: When clean neurons are unlearned via gradient ascent on clean defense data Dd\mathcal{D}_d, the clean classification paths collapse while backdoor pathways remain intact. As a consequence, clean input samples naturally trigger the model's dominant remaining pathways, classifying into the backdoor target class yt=arg⁡max⁡Kf(xd;θ^)y_t = \arg\max_K f(x_d; \hat{\theta}). NU alone achieves 100%100\% top-1 target label detection accuracy across 10 evaluated attack types (BadNets, Trojan, Blend, CL, SIG, Dynamic, WaNet, FC, DFST, AWP), whereas standard Neural Cleanse (NC) achieves only 0%0\% to 95%95\% detection accuracy.
    2. Trigger Pattern Recovery: Running trigger reverse engineering (e.g., NC) on the unlearned model θ^\hat{\theta} rather than the original backdoored model fθf_\theta yields higher fidelity trigger patterns with substantially less background noise, succeeding on stealthy attacks (CL, SIG, Dynamic) where base NC fails.
    3. Backdoor Sample Detection: Combining NU with STRIP (NU+STRIP) widens the entropy difference between clean and backdoored test samples by approximately 0.20.2 relative entropy, improving detection of poisoned samples.
  8. Knowl 8 — Robustness of RNP Against Adaptive Distillation and Adversarial Training Attacks

    empirical result

    RNP was evaluated against two strong adaptive backdoor attacks designed to resist neuron identification on CIFAR-10 with ResNet-18:

    1. Adaptive-Distillation (Adapt-D): Distills the backdoored student network to align intermediate neuron activations with a cleanly trained teacher network.
    2. Adversarial-Training (Adv-T): Perturbs backdoor neurons adversarially during injection and fine-tunes on clean defense data to hide neuron sensitivity.

    Defense results:

    • No Defense: Adapt-D achieves CA 91.56%91.56\%, ASR 84.76%84.76\%; Adv-T achieves CA 88.56%88.56\%, ASR 73.36%73.36\%.
    • Adversarial Neuron Pruning (ANP): Failed against both adaptive attacks, yielding CA 89.78%89.78\%, ASR 24.81%24.81\% on Adapt-D, and CA 83.06%83.06\%, ASR 54.01%54.01\% on Adv-T.
    • RNP (Ours): Maintained robust defense, yielding CA 90.58%90.58\%, ASR 13.22%13.22\% on Adapt-D, and CA 86.15%86.15\%, ASR 18.09%18.09\% on Adv-T.
  9. Knowl 9 — Impact of Defense Dataset Size and Unlearning Epochs on RNP

    empirical result

    Analysis of the hyperparameters governing the Neuron Unlearning (NU) step demonstrates the following behavior:

    1. Defense Dataset Size (∣Dd∣|\mathcal{D}_d|): Evaluating sizes from 0.5%0.5\% (250 samples) to 10%10\% (5000 samples) of the training data shows that ASR reduction is effective across all sizes. However, on complex full-image trigger attacks (CL and Dynamic), increasing defense data to 10%10\% causes Clean Accuracy (CA) to drop steeply from ∼90%\sim 90\% to below 30%30\%. Large clean datasets increase the overlap between unlearned clean features and full-image backdoor representations, causing accidental pruning of clean neurons. Optimal performance occurs at small sample budgets (0.5%0.5\% to 1%1\%).
    2. Unlearning Epochs: Monitoring unlearning progression at 5, 10, 15, and 20 epochs shows that unlearning beyond the point where clean accuracy reaches random guess (CA≈10%CA \approx 10\%, reached within 5-15 epochs) does not improve ASR reduction and causes severe CA drops across attacks (e.g., from gradient explosions during loss maximization). Terminating unlearning immediately when CA reaches random guess prevents model collapse.
  10. Knowl 10 — Limitations of RNP on Lightweight Architectures and Ultra-Low Poisoning Rates

    limitation

    Reconstructive Neuron Pruning (RNP) exhibits three documented limitations:

    1. Lightweight Architectures: On compact models with sparse channel representations like EfficientNet-B0, RNP degrades Clean Accuracy significantly (dropping CA to 63.77%−68.17%63.77\% - 68.17\% on CIFAR-10 across BadNets, Trojan, and Blend) because the limited capacity restricts filter repurposing without damaging essential clean representations.
    2. Ultra-Low Poisoning Rates: When the training poisoning rate drops to ≤0.05%\le 0.05\% (e.g., 250 poisoned images in CIFAR-10 BadNets), backdoor neurons are deeply integrated with clean neurons. RNP achieves 0%0\% ASR but drops CA to 56.32%56.32\% (recovering to 77.29%77.29\% only with post-hoc fine-tuning).
    3. Theoretical Guarantees: RNP is an empirical optimization-based heuristic and does not provide theoretical defense guarantees against arbitrary future attack designs.

Coverage note — None was omitted; all primary conceptual, algorithmic, empirical, ablation, and downstream application contributions are covered.

References

  1. 1.Barni, M., Kallas, K., and Tondi, B. A new backdoor attack in cnns by training set corruption without label poisoning. In ICIP, 2019.
  2. 2.Chen, B., Carvalho, W., Baracaldo, N., Ludwig, H., Edwards, B., Lee, T., Molloy, I., and Srivastava, B. Detecting backdoor attacks on deep neural networks by activation clustering. In AAAI Workshop, 2019.
  3. 3.Chen, W., Wu, B., and Wang, H. Effective backdoor defense by exploiting sensitivity of poisoned samples. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), NeurIPS, 2022.
  4. 4.Chen, X., Liu, C., Li, B., Lu, K., and Song, D. Targeted backdoor attacks on deep learning systems using data poisoning. arXiv preprint arXiv:1712.05526, 2017.
  5. 5.Cheng, S., Liu, Y., Ma, S., and Zhang, X. Deep feature space trojan attack of neural networks by controlled detoxification. In AAAI, 2021.
  6. 6.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  7. 7.Doan, K., Lao, Y., Zhao, W., and Li, P. Lira: Learnable, imperceptible and robust backdoor attacks. In ICCV, 2021.
  8. 8.Gao, Y., Xu, C., Wang, D., Chen, S., Ranasinghe, D. C., and Nepal, S. Strip: A defence against trojan attacks on deep neural networks. In ACSAC, 2019.
  9. 9.Garg, S., Kumar, A., Goel, V., and Liang, Y. Can adversarial weight perturbations inject neural backdoors. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, 2020.
  10. 10.Gu, T., Dolan-Gavitt, B., and Garg, S. Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733, 2017.
  11. 11.Guan, J., Tu, Z., He, R., and Tao, D. Few-shot backdoor defense using shapley estimation. In CVPR, 2022.
  12. 12.Guo, W., Wang, L., Xing, X., Du, M., and Song, D. Tabor: A highly accurate approach to inspecting and restoring trojan backdoors in ai systems. arXiv preprint arXiv:1908.01763, 2019.
  13. 13.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, 2016.
  14. 14.Hu, X., Lin, X., Cogswell, M., Yao, Y., Jha, S., and Chen, C. Trigger hunting with a topological prior for trojan detection. In ICLR, 2022.
  15. 15.Huang, K., Li, Y., Wu, B., Qin, Z., and Ren, K. Backdoor defense via decoupling the training process. In ICLR, 2022.
  16. 16.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  17. 17.Li, Y., Zhai, T., Wu, B., Jiang, Y., Li, Z., and Xia, S. Rethinking the trigger of backdoor attack. arXiv preprint arXiv:2004.04692, 2020.
  18. 18.Li, Y., Li, Y., Wu, B., Li, L., He, R., and Lyu, S. Invisible backdoor attack with sample-specific triggers. In ICCV, pp. 16463–16472, 2021a.
  19. 19.Li, Y., Lyu, X., Koren, N., Lyu, L., Li, B., and Ma, X. Antibackdoor learning: Training clean models on poisoned data. In NeurIPS, 2021b.
  20. 20.Li, Y., Lyu, X., Koren, N., Lyu, L., Li, B., and Ma, X. Neural attention distillation: Erasing backdoor triggers from deep neural networks. In ICLR, 2021c.
  21. 21.Liu, K., Dolan-Gavitt, B., and Garg, S. Fine-pruning: Defending against backdooring attacks on deep neural networks. In RAID, 2018a.
  22. 22.Liu, Y., Ma, S., Aafer, Y., Lee, W.-C., Zhai, J., Wang, W., and Zhang, X. Trojaning attack on neural networks. In NDSS, 2018b.
  23. 23.Liu, Y., Lee, W.-C., Tao, G., Ma, S., Aafer, Y., and Zhang, X. Abs: Scanning neural networks for back-doors by artificial brain stimulation. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security, pp. 1265–1282, 2019.
  24. 24.Liu, Y., Ma, X., Bailey, J., and Lu, F. Reflection backdoor: A natural backdoor attack on deep neural networks. In ECCV, 2020.
  25. 25.Liu, Y., Shen, G., Tao, G., Wang, Z., Ma, S., and Zhang, X. Complex backdoor detection by symmetric feature differencing. In CVPR, 2022.
  26. 26.Nguyen, A. and Tran, A. Input-aware dynamic backdoor attack. In NeurIPS, 2020.
  27. 27.Nguyen, A. and Tran, A. Wanet–imperceptible warpingbased backdoor attack. In ICLR, 2021.
  28. 28.Qi, X., Xie, T., Mahloujifar, S., and Mittal, P. Circumventing backdoor defenses that are based on latent separability. arXiv preprint arXiv:2205.13613, 2022a.
  29. 29.Qi, X., Xie, T., Pan, R., Zhu, J., Yang, Y., and Bu, K. Towards practical deployment-stage backdoor attack on deep neural networks. In CVPR, 2022b.
  30. 30.Shafahi, A., Huang, W. R., Najibi, M., Suciu, O., Studer, C., Dumitras, T., and Goldstein, T. Poison frogs! targeted clean-label poisoning attacks on neural networks. In NeurIPS, 2018.
  31. 31.Shen, G., Liu, Y., Tao, G., An, S., Xu, Q., Cheng, S., Ma, S., and Zhang, X. Backdoor scanning for deep neural networks through k-arm optimization. In ICML, 2021.
  32. 32.Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R. Intriguing properties of neural networks. In ICLR, 2013.
  33. 33.Tang, D., Wang, X., Tang, H., and Zhang, K. Demon in the variant: Statistical analysis of dnns for robust backdoor contamination detection. In USENIX Security, 2021.
  34. 34.Tran, B., Li, J., and Madry, A. Spectral signatures in backdoor attacks. In NeurIPS, 2018.
  35. 35.Turner, A., Tsipras, D., and Madry, A. Clean-label backdoor attacks. https://people.csail.mit.edu/madry/lab/, 2019.
  36. 36.Wang, B., Yao, Y., Shan, S., Li, H., Viswanath, B., Zheng, H., and Zhao, B. Y. Neural cleanse: Identifying and mitigating backdoor attacks in neural networks. In S&P. IEEE, 2019.
  37. 37.Wang, T., Yao, Y., Xu, F., An, S., Tong, H., and Wang, T. An invisible black-box backdoor attack through frequency domain. In ECCV, 2022a.
  38. 38.Wang, Z., Ding, H., Zhai, J., and Ma, S. Training with more confidence: Mitigating injected and natural backdoors during training. In NeurIPS, 2022b.
  39. 39.Wu, D. and Wang, Y. Adversarial neuron pruning purifies backdoored deep models. NeurIPS, 2021.
  40. 40.Xu, X., Wang, Q., Li, H., Borisov, N., Gunter, C. A., and Li, B. Detecting ai trojans using meta neural analysis. In S&P, 2021.
  41. 41.Zeng, Y., Park, W., Mao, Z. M., and Jia, R. Rethinking the backdoor attacks’ triggers: A frequency perspective. In ICCV, 2021.
  42. 42.Zeng, Y., Chen, S., Park, W., Mao, Z. M., Jin, M., and Jia, R. Adversarial unlearning of backdoors via implicit hypergradient. In ICLR, 2022.
  43. 43.Zhao, Z., Chen, X., Xuan, Y., Dong, Y., Wang, D., and Liang, K. Defeat: Deep hidden feature backdoor attacks by imperceptible perturbation and latent representation constraints. In CVPR, 2022.
  44. 44.Zheng, R., Tang, R., Li, J., and Liu, L. Data-free backdoor removal based on channel lipschitzness. In ECCV, 2022.

Citation

MLA
Li, Y., et al. “Reconstructive Neuron Pruning for Backdoor Defense”. International Conference on Machine Learning, vol. 202, 2023, pp. 19837–54, https://proceedings.mlr.press/v202/li23v.html.
APA
Li, Y., Lyu, X., Ma, X., Koren, N., Lyu, L., Li, B., & Jiang, Y.-G. (2023). Reconstructive Neuron Pruning for Backdoor Defense. International Conference on Machine Learning, 202, 19837–19854. https://proceedings.mlr.press/v202/li23v.html
Chicago
Li, Y., X. Lyu, X. Ma, et al. 2023. “Reconstructive Neuron Pruning for Backdoor Defense”. International Conference on Machine Learning 202: 19837–54. https://proceedings.mlr.press/v202/li23v.html.
Harvard
Li, Y. et al. (2023) “Reconstructive Neuron Pruning for Backdoor Defense”, International Conference on Machine Learning. PMLR, pp. 19837–19854. Available at: https://proceedings.mlr.press/v202/li23v.html.
Vancouver
1. Li Y, Lyu X, Ma X, Koren N, Lyu L, Li B, Jiang Y-G (2023) Reconstructive Neuron Pruning for Backdoor Defense. In: International Conference on Machine Learning. PMLR, pp 19837–19854

BibTeX

@InProceedings{pmlr-v202-li23v,
  title = 	 {Reconstructive Neuron Pruning for Backdoor Defense},
  author =       {Li, Yige and Lyu, Xixiang and Ma, Xingjun and Koren, Nodens and Lyu, Lingjuan and Li, Bo and Jiang, Yu-Gang},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {19837--19854},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/li23v/li23v.pdf},
  url = 	 {https://proceedings.mlr.press/v202/li23v.html},
  abstract = 	 {Deep neural networks (DNNs) have been found to be vulnerable to backdoor attacks, raising security concerns about their deployment in mission-critical applications. While existing defense methods have demonstrated promising results, it is still not clear how to effectively remove backdoor-associated neurons in backdoored DNNs. In this paper, we propose a novel defense called Reconstructive Neuron Pruning (RNP) to expose and prune backdoor neurons via an unlearning and then recovering process. Specifically, RNP first unlearns the neurons by maximizing the model’s error on a small subset of clean samples and then recovers the neurons by minimizing the model’s error on the same data. In RNP, unlearning is operated at the neuron level while recovering is operated at the filter level, forming an asymmetric reconstructive learning procedure. We show that such an asymmetric process on only a few clean samples can effectively expose and prune the backdoor neurons implanted by a wide range of attacks, achieving a new state-of-the-art defense performance. Moreover, the unlearned model at the intermediate step of our RNP can be directly used to improve other backdoor defense tasks including backdoor removal, trigger recovery, backdoor label detection, and backdoor sample detection. Code is available at https://github.com/bboylyg/RNP.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/