Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks

Kang LiuBrendan Dolan-GavittSiddharth Garg

article2018International Symposium on Recent Advances in Intrusion Detection1,386 citations

Proposes Fine-Pruning, a defense combining neural network pruning and fine-tuning that neutralizes backdoor attacks from untrusted training pipelines while maintaining classification accuracy on clean inputs.

Listen

Modern organizations frequently outsource the computationally demanding training of deep learning models to third-party cloud providers. However, this practice creates severe security risks: a malicious provider can plant a hidden backdoor inside the model. Under normal circumstances, a backdoored model performs accurately and passes validation checks, but when presented with an attacker's specific trigger—such as an image pattern or audio noise—it triggers targeted misclassifications or systemic failures. Traditional validation data cannot detect these dormant vulnerabilities, exposing deployed systems to critical operational and safety hazards.

The article aims to design and evaluate effective defensive strategies capable of neutralizing backdoor attacks in outsourced deep neural networks without sacrificing baseline model performance. Specifically, the article demonstrates the vulnerability of standard defenses and evaluates a combined mitigation strategy across multiple real-world vision and speech applications.

To assess these defenses, the article reproduces three known backdoor attacks across face recognition, speech recognition, and traffic sign detection models. The investigation tests two baseline countermeasures: pruning, which removes dormant neurons from the network, and fine-tuning, which briefly retrains the model using a small set of clean data. The authors also construct a sophisticated "pruning-aware" attack designed to evade detection by forcing clean and backdoored behaviors to share the same neurons. Finally, the article evaluates "fine-pruning," a two-step defense combining neuron pruning with subsequent clean fine-tuning.

The findings demonstrate clear limitations in single-defense strategies and prove the efficacy of the combined approach. First, while simple pruning removes backdoors from standard attacks, the advanced pruning-aware attack completely evades it, causing clean classification accuracy to collapse before the backdoor is disabled. Second, fine-tuning alone fails against baseline attacks because backdoor neurons remain dormant during clean retraining and do not update their parameters. Third, fine-pruning successfully neutralizes both standard and sophisticated attacks across all domains. For targeted attacks, fine-pruning drops backdoor success rates to 0% in face recognition and down to 0% to 2% in speech recognition, while clean data accuracy remains essentially unchanged (falling by at most 0.2%). For untargeted traffic sign attacks, fine-pruning reduces backdoor attack success from 90%–99% down to 29%–37%.

These results provide a practical path for safely adopting third-party machine learning services. Organizations do not need to choose between the prohibitive expense of full in-house training and the security risks of untrusted providers. By requiring only a brief local post-processing step that converges in minutes, fine-pruning eliminates backdoors at a fraction of the computational and financial cost of training from scratch.

Organizations that deploy outsourced deep learning models should implement fine-pruning as a standard security control prior to deployment. Game-theoretic analysis in the article confirms that fine-pruning is the optimal defensive strategy regardless of the attacker's approach. When configuring the defense, teams should prune dormant neurons until clean validation accuracy declines slightly, followed by fine-tuning on a clean dataset until convergence to restore accuracy and eliminate residual triggers.

Confidence in these findings is high for convolutional architectures using standard activation functions, as the defense was validated across multiple distinct domains. However, the article notes a key limitation: the evaluations focus on convolutional neural networks, meaning effectiveness on other architectures, such as recurrent networks used in sequential language tasks, remains unverified and requires further study.

  • Paper: How To Backdoor Federated Learning, Eugene Bagdasaryan et al. (2018). This work extends backdoor injection techniques from centralized outsourcing settings to federated learning environments where centralized defensive pruning is harder to apply.
  • Paper: Rethinking the Value of Network Pruning, Zhuang Liu et al. (2019). This study critically re-evaluates the role of pruning and fine-tuning in neural networks, challenging the fundamental assumptions about weight retention that underlie fine-pruning pipelines.
  • Paper: The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks., Jonathan Frankle et al. (2019). It deepens the understanding of pruning mechanics by demonstrating the existence of trainable sparse subnetworks discovered through iterative magnitude pruning.
Cover for Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks

Abstract

Deep neural networks (DNNs) provide excellent performance across a wide range of classification tasks, but their training requires high computational resources and is often outsourced to third parties. Recent work has shown that outsourced training introduces the risk that a malicious trainer will return a backdoored DNN that behaves normally on most inputs but causes targeted misclassifications or degrades the accuracy of the network when a trigger known only to the attacker is present. In this paper, we provide the first effective defenses against backdoor attacks on DNNs. We implement three backdoor attacks from prior work and use them to investigate two promising defenses, pruning and fine-tuning. We show that neither, by itself, is sufficient to defend against sophisticated attackers. We then evaluate fine-pruning, a combination of pruning and fine-tuning, and show that it successfully weakens or even eliminates the backdoors, i.e., in some cases reducing the attack success rate to 0% with only a 0.4% drop in accuracy for clean (non-triggering) inputs. Our work provides the first step toward defenses against backdoor attacks in deep neural networks.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Neural Network Basics
  • Deep Neural Networks (DNN)
  • DNN Training
  • 2.2 Threat Model
  • Setting
  • Attacker’s Goals
  • Attacker’s Capabilities
  • 2.3 Backdoor Attacks
  • Face Recognition Backdoor
  • Speech Recognition Backdoor
  • Traffic Sign Backdoor
  • 3 Methodology
  • 3.1 Pruning Defense
  • 3.2 Pruning-Aware Attack
  • 3.3 Fine-Pruning Defense
  • 4 Discussion
  • 4.1 Threats to Validity
  • 5 Related Work
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Fine-Pruning Defense Against Neural Network Backdoors

    algorithm

    The Fine-Pruning defense removes backdoor vulnerabilities from a trained deep neural network by first pruning dormant neurons and then fine-tuning the remaining parameters on a small, clean validation dataset.

    Input: Backdoored model FΘF_{\Theta}, clean validation dataset DvalidD_{valid}, target convolutional layer ll, clean accuracy drop threshold τ\tau, learning rate η\eta, fine-tuning epochs EE
    Output: Defended model FΘ∗F_{\Theta^*}
    for each filter/neuron jj in layer ll do
        aj←1∣Dvalid∣∑x∈DvalidActivationj(FΘ,x)a_j \leftarrow \frac{1}{|D_{valid}|} \sum_{x \in D_{valid}} \text{Activation}_j(F_{\Theta}, x)
    end for
    Sort neurons in layer ll in ascending order of mean clean activation: (j1,j2,…,jNl)(j_1, j_2, \dots, j_{N_l})
    for m←1m \leftarrow 1 to NlN_l do
        Prune neuron jmj_m from layer ll of FΘF_{\Theta}
        if Accuracy(FΘ,Dvalid)\text{Accuracy}(F_{\Theta}, D_{valid}) drops by more than τ\tau from original accuracy then
            break
        end if
    end for
    Initialize fine-tuning optimizer with current pruned parameters Θ\Theta
    for epoch←1\text{epoch} \leftarrow 1 to EE do
        for each batch (X,Z)⊆Dvalid(X, Z) \subseteq D_{valid} do
            Θ←Θ−η∇ΘL(FΘ(X),Z)\Theta \leftarrow \Theta - \eta \nabla_{\Theta} \mathcal{L}(F_{\Theta}(X), Z)
        end for
    end for
    return FΘF_{\Theta}

    Fine-pruning overcomes the limitations of using either pruning or fine-tuning in isolation. Against standard backdoor attacks where backdoor neurons remain dormant on clean data, the pruning phase removes those backdoor neurons directly, and subsequent fine-tuning restores the slight classification accuracy loss on clean inputs. Against pruning-aware backdoor attacks where the backdoor is co-located with clean features on active neurons, pruning removes unneeded decoy neurons, and subsequent fine-tuning on clean data updates the shared weights, forcing the network to unlearn the backdoor trigger.

  2. Knowl 2 — Pruning-Aware Backdoor Attack

    algorithm

    The pruning-aware backdoor attack embeds a backdoor into a deep neural network while evading activation-based pruning defenses by concentrating both clean classification and backdoor functionality onto a shared, compact subset of neurons, while re-inserting dormant 'decoy' neurons into the network architecture.

    Input: Baseline architecture FF, clean dataset DcleanD_{clean}, poisoned dataset DpoisonD_{poison}, target prune ratio pp, bias reduction Δb\Delta b
    Output: Backdoored model parameters Θ′\Theta'
    Train FΘF_{\Theta} on DcleanD_{clean} until convergence
    Compute mean activation on DcleanD_{clean} for all neurons in the target convolutional layer
    Prune the lowest fraction pp of dormant neurons to create reduced architecture FprunedF_{pruned}
    Train FprunedF_{pruned} on DpoisonD_{poison} until clean accuracy and attack success reach desired thresholds
    if capacity is insufficient then
        Reinstate individual neurons into FprunedF_{pruned} and resume training
    end if
    Reinstate all remaining pruned neurons into FprunedF_{pruned} with original pre-trained weights to match architecture FF
    for each reinstated neuron kk do
        bk←bk−Δbb_k \leftarrow b_k - \Delta b to ensure activation on DcleanD_{clean} is strictly zero
    end for
    return Model parameters Θ′\Theta'

    When a defender applies pruning based on clean input activations, the defender's procedure selects and prunes the reinstated decoy neurons first, leaving the unpruned core carrying the backdoor untouched.

  3. Knowl 3 — Pruning Defense Against Spare-Capacity Backdoor Attacks

    model/method

    Backdoor attacks on deep neural networks exploit spare model capacity by allocating backdoor features to neurons that remain dormant during inference on clean inputs. The pruning defense identifies and disables these backdoor channels without knowledge of the backdoor trigger.

    For each neuron jj in the final convolutional layer of a network FΘF_\Theta, the defender computes the average activation over a held-out clean validation set DvalidD_{valid}:

    aˉj=1∣Dvalid∣∑x∈Dvalidaj(x)\bar{a}_j = \frac{1}{|D_{valid}|} \sum_{x \in D_{valid}} a_j(x)

    Neurons are iteratively removed in ascending order of aˉj\bar{a}_j. Pruning proceeds through three characteristic phases:

    1. Dormant neurons that activate on neither clean nor backdoored inputs are pruned with no change to clean accuracy or attack success.
    2. Backdoor-dedicated neurons that activate only on triggered inputs are pruned, causing a sharp drop in backdoor attack success rate while maintaining high clean accuracy.
    3. Clean-feature neurons begin to be pruned, causing clean validation accuracy to degrade.

    Terminating pruning when clean validation accuracy drops by ≤4%\le 4\% reduces backdoor attack success from 99%99\% to 0%0\% on DeepID face recognition, from 77%77\% to 13%13\% on AlexNet speech recognition, and from 98%98\% to 35%35\% on Faster-RCNN traffic sign detection.

  4. Knowl 4 — Empirical Comparison of Fine-Tuning and Fine-Pruning Against Backdoor Attacks

    data/table

    The following table reports classification accuracy on clean inputs (clcl) and backdoor attack success rate (bdbd) under No Defense, Fine-Tuning, and Fine-Pruning against baseline and pruning-aware backdoor attacks across face recognition (DeepID), speech recognition (AlexNet), and traffic sign detection (Faster-RCNN):

    Neural Network Baseline Attack Pruning-Aware Attack
    None Fine-Tuning Fine-Pruning None Fine-Tuning Fine-Pruning
    Face Recognition cl: 0.978 cl: 0.978 cl: 0.978 cl: 0.974 cl: 0.978 cl: 0.977
    bd: 1.000 bd: 0.000 bd: 0.000 bd: 0.998 bd: 0.000 bd: 0.000
    Speech Recognition cl: 0.990 cl: 0.990 cl: 0.988 cl: 0.988 cl: 0.988 cl: 0.986
    bd: 0.770 bd: 0.435 bd: 0.020 bd: 0.780 bd: 0.520 bd: 0.000
    Traffic Sign Detection cl: 0.849 cl: 0.857 cl: 0.873 cl: 0.820 cl: 0.872 cl: 0.874
    bd: 0.991 bd: 0.921 bd: 0.288 bd: 0.899 bd: 0.419 bd: 0.366

    Fine-pruning consistently outperforms standalone fine-tuning. For targeted attacks (face and speech recognition), fine-pruning reduces attack success to 0.0%–2.0%0.0\%\text{--}2.0\% with at most a 0.2%0.2\% drop in clean accuracy (and in some cases slight accuracy improvements). For untargeted traffic sign detection, fine-pruning reduces attack success from 99.1%99.1\% down to 28.8%28.8\% in the baseline attack and from 89.9%89.9\% down to 36.6%36.6\% in the pruning-aware attack, whereas fine-tuning alone leaves baseline attack success at 92.1%92.1\%.

  5. Knowl 5 — Game-Theoretic Utility Matrix of Backdoor Defense Strategies

    data/table

    In an adversarial setting modeled as a zero-sum game between an attacker choosing an attack strategy and a defender choosing a defense strategy, the defender's utility is defined as U=Accuracyclean−SuccessbackdoorU = \text{Accuracy}_{clean} - \text{Success}_{backdoor}. The utility matrix evaluated on speech recognition is:

    Defender Strategy Baseline Attack Pruning-Aware Attack
    Fine-Tuning 0.555 0.468
    Fine-Pruning 0.968 0.986

    Fine-pruning is a strictly dominant strategy for the defender. If the defender employs only fine-tuning, the attacker achieves higher payoff with a baseline attack (0.5550.555 defender utility vs 0.9680.968 under fine-pruning). Employing fine-pruning guarantees a defender utility ≥0.968\ge 0.968 regardless of the attacker's chosen strategy.

  6. Knowl 6 — Failure Mechanism of Standalone Fine-Tuning on Dormant Backdoor Neurons

    theoretical result

    In a neural network trained with a baseline backdoor attack, neurons encoding the backdoor trigger remain dormant (activation aj=0a_j = 0 under ReLU) on all clean training inputs. When fine-tuning is conducted using a clean dataset DvalidD_{valid}, the pre-activation input zj=wjaj−1+bjz_j = w_j a_{j-1} + b_j is negative for these neurons, yielding an activation derivative of zero:

    ∂aj∂zj=0\frac{\partial a_j}{\partial z_j} = 0

    Consequently, the gradient of the clean training loss L\mathcal{L} with respect to the weights wjw_j and bias bjb_j of the backdoor neurons evaluates to zero:

    ∇wjL=∂L∂aj∂aj∂zjaj−1=0\nabla_{w_j} \mathcal{L} = \frac{\partial \mathcal{L}}{\partial a_j} \frac{\partial a_j}{\partial z_j} a_{j-1} = 0

    Standard gradient descent updates Θ←Θ−η∇ΘL\Theta \leftarrow \Theta - \eta \nabla_\Theta \mathcal{L} therefore leave the weights of dormant backdoor neurons unchanged, allowing the backdoor trigger to persist through fine-tuning.

  7. Knowl 7 — Evasion of Pruning Defenses via Decoy Neurons in Pruning-Aware Attacks

    empirical result

    The pruning-aware attack evades activation-based pruning by concentrating clean and backdoor activations onto a shared 3%–16%3\%\text{--}16\% of neurons in the final convolutional layer, while converting the remaining neurons into dormant 'decoys' via reduced biases.

    When a defender prunes neurons in ascending order of clean activations, all decoy neurons are pruned first without affecting either clean accuracy or backdoor success. Once pruning reaches active neurons, the defense degrades clean classification accuracy before disabling backdoor behavior:

    • In face recognition, clean accuracy falls below 23%23\% before any reduction in backdoor success rate is observed.
    • In speech recognition, clean accuracy drops by 55%55\% to achieve the same reduction in backdoor success (13%13\%) that required only a 4%4\% accuracy drop against the baseline attack.
    • In traffic sign recognition, attack success remains high across all pruning levels until clean accuracy is destroyed.
  8. Knowl 8 — Threat Model for Outsourced Neural Network Training Under Backdoor Attacks

    assumption

    The threat model considers deep neural network training outsourced to an untrusted third party:

    1. User Resources and Goals: The user supplies a clean training dataset DtrainD_{train} and model architecture specification FF to the third party. The user retains a private, clean validation dataset DvalidD_{valid} unavailable to the third party. The user accepts and deploys the returned model parameters Θ′\Theta' only if validation accuracy on DvalidD_{valid} exceeds a predefined threshold.

    2. Attacker Capabilities: The attacker has white-box access to the training pipeline, including arbitrary modification or poisoning of DtrainD_{train}, adjustment of training hyperparameters (epochs, learning rate, batch size), or manual setting of network weights Θ′\Theta', but cannot alter the user-specified network architecture or hyper-parameters.

    3. Attacker Goals:

    • Trigger-Activated Misbehavior: For any input xx containing an attacker-specified trigger, the model outputs an incorrect target class (targeted attack) or arbitrary incorrect class (untargeted attack).
    • Validation Stealth: The model maintains accuracy on clean validation inputs comparable to an honestly trained model so that it passes user validation.
  9. Knowl 9 — Attack Success Rate Metric for Untargeted Backdoors

    equation

    For untargeted backdoor attacks where the adversarial objective is to cause the model to misclassify any triggered input into an arbitrary incorrect class, the attack success rate ASRuntargetedASR_{untargeted} is defined as:

    ASRuntargeted=1−AbackdoorAcleanASR_{untargeted} = 1 - \frac{A_{backdoor}}{A_{clean}}

    where Aclean∈[0,1]A_{clean} \in [0, 1] denotes the classification accuracy of the network evaluated on clean inputs, and Abackdoor∈[0,1]A_{backdoor} \in [0, 1] denotes the classification accuracy evaluated on backdoored inputs containing the trigger pattern. When the trigger causes complete misclassification (Abackdoor=0A_{backdoor} = 0), ASRuntargeted=1.0ASR_{untargeted} = 1.0; when the trigger has no impact on accuracy (Abackdoor=AcleanA_{backdoor} = A_{clean}), ASRuntargeted=0ASR_{untargeted} = 0.

  10. Knowl 10 — Limitations of the Fine-Pruning Defense

    limitation

    The fine-pruning defense exhibits three key limitations:

    1. Architectural Scope: Fine-pruning was evaluated on Convolutional Neural Networks (CNNs) employing ReLU activation functions. Its efficacy on sequential or recurrent architectures such as Recurrent Neural Networks (RNNs) and Long Short-Term Memory networks (LSTMs) remains unverified.
    2. Potential Adversarial Initialization: Fine-tuning continues gradient descent from the attacker-provided parameter state Θi\Theta_i. An attacker who identifies an initialization Θi\Theta_i residing in a local minimum on the clean loss surface that retains the backdoor could evade fine-pruning unless Gaussian noise is injected into parameters prior to fine-tuning.
    3. Residual Vulnerability in Untargeted Settings: While fine-pruning reduces targeted backdoor success to near zero (0.0%–2.0%0.0\%\text{--}2.0\%), in untargeted settings (Faster-RCNN traffic sign detection) the attack success rate after fine-pruning remains between 28.8%28.8\% and 36.6%36.6\% because any misclassification on a triggered input satisfies the untargeted attack objective.

Coverage note — No substantial contributed material was omitted; all key contributions, including baseline attack replications, the pruning defense, the pruning-aware attack, the fine-pruning algorithm, empirical evaluations, utility matrix analysis, and defense limitations, are covered.

References

  1. 1.ImageNet large scale visual recognition competition. http://www.image-net.org/challenges/LSVRC/2012/, 2012.
  2. 2.Amazon Web Services, Inc. Amazon Elastic Compute Cloud (Amazon EC2). https://aws.amazon.com/ec2/.
  3. 3.Amazon.com, Inc. Deep Learning AMI Amazon Linux Version.
  4. 4.S. Anwar et al. Structured pruning of deep convolutional neural networks. ACM Journal on Emerging Technologies in Computing Systems (JETC), 13(3):32, 2017.
  5. 5.A. Athalye, N. Carlini, and D. Wagner. Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples. ArXiv e-prints, Feb. 2018.
  6. 6.D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly learning to align and translate, 2014.
  7. 7.M. Barreno, B. Nelson, R. Sears, A. D. Joseph, and J. D. Tygar. Can machine learning be secure? In Proceedings of the 2006 ACM Symposium on Information, Computer and Communications Security, ASIACCS ’06, pages 16–25, New York, NY, USA, 2006. ACM.
  8. 8.A. Blum and R. L. Rivest. Training a 3-node neural network is np-complete. In Advances in neural information processing systems, pages 494–501, 1989.
  9. 9.N. Carlini and D. A. Wagner. Defensive distillation is not robust to adversarial examples. CoRR, abs/1607.04311, 2016.
  10. 10.X. Chen, C. Liu, B. Li, K. Lu, and D. Song. Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning. ArXiv e-prints, Dec. 2017.
  11. 11.S. P. Chung and A. K. Mok. Allergy attack against automatic signature generation. In Proceedings of the 9th International Conference on Recent Advances in Intrusion Detection, 2006.
  12. 12.S. P. Chung and A. K. Mok. Advanced allergy attacks: Does a corpus really help. In Proceedings of the 10th International Conference on Recent Advances in Intrusion Detection, 2007.
  13. 13.G. S. Dhillon, K. Azizzadenesheli, J. D. Bernstein, J. Kossaifi, A. Khanna, Z. C. Lipton, and A. Anandkumar. Stochastic activation pruning for robust adversarial defense. In International Conference on Learning Representations, 2018.
  14. 14.P. Fogla and W. Lee. Evading network anomaly detection systems: Formal reasoning and practical techniques. In Proceedings of the 13th ACM Conference on Computer and Communications Security, CCS ’06, pages 59–68, New York, NY, USA, 2006. ACM.
  15. 15.P. Fogla, M. Sharif, R. Perdisci, O. Kolesnikov, and W. Lee. Polymorphic blending attacks. In Proceedings of the 15th Conference on USENIX Security Symposium - Volume 15, USENIX-SS’06, Berkeley, CA, USA, 2006. USENIX Association.
  16. 16.Google, Inc. Google Cloud Machine Learning Engine. https://cloud.google.com/ml-engine/.
  17. 17.A. Graves, A.-r. Mohamed, and G. Hinton. Speech recognition with deep recurrent neural networks. In Acoustics, speech and signal processing (icassp), 2013 ieee international conference on, pages 6645–6649. IEEE, 2013.
  18. 18.T. Gu, S. Garg, and B. Dolan-Gavitt. BadNets: Identifying vulnerabilities in the machine learning model supply chain. In NIPS Machine Learning and Computer Security Workshop, 2017. https://arxiv.org/abs/1708.06733.
  19. 19.S. Han et al. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149, 2015.
  20. 20.W. He, J. Wei, X. Chen, N. Carlini, and D. Song. Adversarial example defense: Ensembles of weak defenses are not strong. In 11th USENIX Workshop on Offensive Technologies (WOOT 17), Vancouver, BC, 2017. USENIX Association.
  21. 21.K. M. Hermann and P. Blunsom. Multilingual Distributed Representations without Word Alignment. In Proceedings of ICLR, Apr. 2014.
  22. 22.F. N. Iandola, M. W. Moskewicz, K. Ashraf, and K. Keutzer. Firecaffe: near-linear acceleration of deep neural network training on compute clusters. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2592–2600, 2016.
  23. 23.C. Karlberger, G. Bayler, C. Kruegel, and E. Kirda. Exploiting redundancy in natural language to penetrate bayesian spam filters. In Proceedings of the First USENIX Workshop on Offensive Technologies, WOOT ’07, pages 9:1–9:7, Berkeley, CA, USA, 2007. USENIX Association.
  24. 24.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012.
  25. 25.H. Li et al. Pruning filters for efficient convnets. arXiv preprint arXiv:1608.08710, 2016.
  26. 26.C. Liu, B. Li, Y. Vorobeychik, and A. Oprea. Robust linear regression against training data poisoning. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, pages 91–102. ACM, 2017.
  27. 27.Y. Liu, S. Ma, Y. Aafer, W.-C. Lee, J. Zhai, W. Wang, and X. Zhang. Trojaning attack on neural networks. In 25nd Annual Network and Distributed System Security Symposium, NDSS 2018, San Diego, California, USA, February 18-221, 2018. The Internet Society, 2018.
  28. 28.Y. Liu, Y. Xie, and A. Srivastava. Neural trojans. CoRR, abs/1710.00942, 2017.
  29. 29.D. Lowd and C. Meek. Adversarial learning. In Proceedings of the Eleventh ACM SIGKDD International Conference on Knowledge Discovery in Data Mining, KDD ’05, pages 641–647, New York, NY, USA, 2005. ACM.
  30. 30.D. Lowd and C. Meek. Good word attacks on statistical spam filters. In Proceedings of the Conference on Email and Anti-Spam (CEAS), 2005.
  31. 31.Microsoft Corp. Azure Batch AI Training. https://batchaitraining.azure.com/.
  32. 32.A. Møgelmose, D. Liu, and M. M. Trivedi. Traffic sign detection for us roads: Remaining challenges and a case for tracking. In Intelligent Transportation Systems (ITSC), 2014 IEEE 17th International Conference on, pages 1394–1399. IEEE, 2014.
  33. 33.P. Molchanov et al. Pruning convolutional neural networks for resource efficient inference. 2016.
  34. 34.L. Muñoz-González, B. Biggio, A. Demontis, A. Paudice, V. Wongrassamee, E. C. Lupu, and F. Roli. Towards poisoning of deep learning algorithms with back-gradient optimization. CoRR, abs/1708.08689, 2017.
  35. 35.B. Nelson, M. Barreno, F. J. Chi, A. D. Joseph, B. I. P. Rubinstein, U. Saini, C. Sutton, J. D. Tygar, and K. Xia. Exploiting machine learning to subvert your spam filter. In Proceedings of the 1st Usenix Workshop on Large-Scale Exploits and Emergent Threats, LEET’08, pages 7:1–7:9, Berkeley, CA, USA, 2008. USENIX Association.
  36. 36.J. Newsome, B. Karp, and D. Song. Paragraph: Thwarting signature learning by training maliciously. In Proceedings of the 9th International Conference on Recent Advances in Intrusion Detection, RAID’06, pages 81–105, Berlin, Heidelberg, 2006. Springer-Verlag.
  37. 37.N. Papernot, P. McDaniel, X. Wu, S. Jha, and A. Swami. Distillation as a defense to adversarial perturbations against deep neural networks. In 2016 IEEE Symposium on Security and Privacy (SP), pages 582–597, May 2016.
  38. 38.S. Ren, K. He, R. Girshick, and J. Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems, pages 91–99, 2015.
  39. 39.O. Suciu, R. Mărginean, Y. Kaya, H. Daumé III, and T. Dumitraș. When does machine learning fail? generalized transferability for evasion and poisoning attacks. arXiv preprint arXiv:1803.06975, 2018.
  40. 40.Y. Sun, X. Wang, and X. Tang. Deep learning face representation from predicting 10,000 classes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1891–1898, 2014.
  41. 41.K. M. C. Tan, K. S. Killourhy, and R. A. Maxion. Undermining an anomaly-based intrusion detection system using common exploits. In Proceedings of the 5th International Conference on Recent Advances in Intrusion Detection, RAID’02, pages 54–73, Berlin, Heidelberg, 2002. Springer-Verlag.
  42. 42.F. Tung, S. Muralidharan, and G. Mori. Fine-pruning: Joint fine-tuning and compression of a convolutional network with bayesian optimization. arXiv preprint arXiv:1707.09102, 2017.
  43. 43.D. Wagner and P. Soto. Mimicry attacks on host-based intrusion detection systems. In Proceedings of the 9th ACM Conference on Computer and Communications Security, CCS ’02, pages 255–264, New York, NY, USA, 2002. ACM.
  44. 44.G. L. Wittel and S. F. Wu. On Attacking Statistical Spam Filters. In Proceedings of the Conference on Email and Anti-Spam (CEAS), Mountain View, CA, USA, 2004.
  45. 45.L. Wolf, T. Hassner, and I. Maoz. Face recognition in unconstrained videos with matched background similarity. In CVPR 2011, pages 529–534, June 2011.
  46. 46.H. Xiao, K. Rasul, and R. Vollgraf. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. CoRR, abs/1708.07747, 2017.
  47. 47.J. Yosinski, J. Clune, Y. Bengio, and H. Lipson. How transferable are features in deep neural networks? In Advances in neural information processing systems, pages 3320–3328, 2014.
  48. 48.J. Yu et al. Scalpel: Customizing dnn pruning to the underlying hardware parallelism. In Proceedings of the 44th Annual International Symposium on Computer Architecture, pages 548–560. ACM, 2017.

Citation

MLA
Liu, K., et al. “Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks”. arXiv, 2018, http://arxiv.org/abs/1805.12185v1.
APA
Liu, K., Dolan-Gavitt, B., & Garg, S. (2018). Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks. arXiv. http://arxiv.org/abs/1805.12185v1
Chicago
Liu, K., B. Dolan-Gavitt, and S. Garg. 2018. “Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks”. arXiv. http://arxiv.org/abs/1805.12185v1.
Harvard
Liu, K., Dolan-Gavitt, B. and Garg, S. (2018) “Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1805.12185v1.
Vancouver
1. Liu K, Dolan-Gavitt B, Garg S (2018) Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks. arXiv

BibTeX

@article{liu2018fine,
  title = {Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks},
  author = {Liu, Kang and Dolan-Gavitt, Brendan and Garg, Siddharth},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1805.12185v1},
  eprint = {1805.12185}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF