Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks
Kang LiuBrendan Dolan-GavittSiddharth Garg
Proposes Fine-Pruning, a defense combining neural network pruning and fine-tuning that neutralizes backdoor attacks from untrusted training pipelines while maintaining classification accuracy on clean inputs.
Modern organizations frequently outsource the computationally demanding training of deep learning models to third-party cloud providers. However, this practice creates severe security risks: a malicious provider can plant a hidden backdoor inside the model. Under normal circumstances, a backdoored model performs accurately and passes validation checks, but when presented with an attacker's specific trigger—such as an image pattern or audio noise—it triggers targeted misclassifications or systemic failures. Traditional validation data cannot detect these dormant vulnerabilities, exposing deployed systems to critical operational and safety hazards.
The article aims to design and evaluate effective defensive strategies capable of neutralizing backdoor attacks in outsourced deep neural networks without sacrificing baseline model performance. Specifically, the article demonstrates the vulnerability of standard defenses and evaluates a combined mitigation strategy across multiple real-world vision and speech applications.
To assess these defenses, the article reproduces three known backdoor attacks across face recognition, speech recognition, and traffic sign detection models. The investigation tests two baseline countermeasures: pruning, which removes dormant neurons from the network, and fine-tuning, which briefly retrains the model using a small set of clean data. The authors also construct a sophisticated "pruning-aware" attack designed to evade detection by forcing clean and backdoored behaviors to share the same neurons. Finally, the article evaluates "fine-pruning," a two-step defense combining neuron pruning with subsequent clean fine-tuning.
The findings demonstrate clear limitations in single-defense strategies and prove the efficacy of the combined approach. First, while simple pruning removes backdoors from standard attacks, the advanced pruning-aware attack completely evades it, causing clean classification accuracy to collapse before the backdoor is disabled. Second, fine-tuning alone fails against baseline attacks because backdoor neurons remain dormant during clean retraining and do not update their parameters. Third, fine-pruning successfully neutralizes both standard and sophisticated attacks across all domains. For targeted attacks, fine-pruning drops backdoor success rates to 0% in face recognition and down to 0% to 2% in speech recognition, while clean data accuracy remains essentially unchanged (falling by at most 0.2%). For untargeted traffic sign attacks, fine-pruning reduces backdoor attack success from 90%–99% down to 29%–37%.
These results provide a practical path for safely adopting third-party machine learning services. Organizations do not need to choose between the prohibitive expense of full in-house training and the security risks of untrusted providers. By requiring only a brief local post-processing step that converges in minutes, fine-pruning eliminates backdoors at a fraction of the computational and financial cost of training from scratch.
Organizations that deploy outsourced deep learning models should implement fine-pruning as a standard security control prior to deployment. Game-theoretic analysis in the article confirms that fine-pruning is the optimal defensive strategy regardless of the attacker's approach. When configuring the defense, teams should prune dormant neurons until clean validation accuracy declines slightly, followed by fine-tuning on a clean dataset until convergence to restore accuracy and eliminate residual triggers.
Confidence in these findings is high for convolutional architectures using standard activation functions, as the defense was validated across multiple distinct domains. However, the article notes a key limitation: the evaluations focus on convolutional neural networks, meaning effectiveness on other architectures, such as recurrent networks used in sequential language tasks, remains unverified and requires further study.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). This foundational paper establishes the BadNets backdoor attack paradigm on neural networks that Fine-Pruning directly evaluates and seeks to defend against.
- Paper: Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning, Xinyun Chen et al. (2017). It introduces targeted data poisoning backdoor attacks on deep learning architectures, providing the specific attack methodologies mitigated by the Fine-Pruning defense.
- Paper: Learning both Weights and Connections for Efficient Neural Networks, Song Han et al. (2015). It introduces the classic pipeline of pruning network weights followed by fine-tuning, which forms the underlying architectural mechanics adapted for backdoor defense.
- Paper: Learning Efficient Convolutional Networks through Network Slimming, Zhuang Liu et al. (2017). It formalizes channel-level network pruning and subsequent fine-tuning to compress deep neural networks without compromising baseline classification accuracy.
- Paper: Poisoning Attacks against Support Vector Machines, Battista Biggio et al. (2012). It provides foundational optimization techniques for data poisoning attacks against learning systems that motivate training-time security defenses.
- Paper: How To Backdoor Federated Learning, Eugene Bagdasaryan et al. (2018). This work extends backdoor injection techniques from centralized outsourcing settings to federated learning environments where centralized defensive pruning is harder to apply.
- Paper: Rethinking the Value of Network Pruning, Zhuang Liu et al. (2019). This study critically re-evaluates the role of pruning and fine-tuning in neural networks, challenging the fundamental assumptions about weight retention that underlie fine-pruning pipelines.
- Paper: The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks., Jonathan Frankle et al. (2019). It deepens the understanding of pruning mechanics by demonstrating the existence of trainable sparse subnetworks discovered through iterative magnitude pruning.
