Poisoning Attacks against Support Vector Machines
Battista BiggioBlaine NelsonPavel Laskov
Proposes a gradient-based poisoning attack that generates malicious training data to maximize validation error, demonstrating how intelligent adversaries can systematically subvert Support Vector Machines across linear and non-linear kernels.
Machine learning systems are widely deployed in security-critical environments, such as spam filtering, malware classification, and intrusion detection. However, standard learning models implicitly assume that training data originates from a benign, well-behaved distribution. In adversarial contexts, attackers can manipulate the data fed into these systems. The article addresses the emerging threat of data poisoning, a scenario where malicious actors deliberately inject corrupted samples into training sets to degrade classifier performance.
The objective of the article is to formulate and demonstrate an optimization framework that generates highly effective malicious training points to systematically degrade the classification accuracy of Support Vector Machines (SVMs), a standard machine learning model.
To achieve this, the authors developed a gradient ascent algorithm that computes how an injected sample shifts the optimal decision boundary of an SVM. Unlike previous approaches that could only design attacks within abstract internal feature representations, this method operates directly in the raw input space across linear and non-linear configurations. The attack was tested on synthetic two-dimensional datasets as well as real-world image classification benchmarks (using the MNIST handwritten digit repository) by establishing a baseline model, selecting an initial mislabeled point, and iteratively modifying it to maximize classification error.
The experimental findings show that deliberate data poisoning is exceptionally potent. First, the gradient-based optimization successfully identified local maxima on the error surface, causing far greater disruption than random label flipping or random noise. Second, injecting just a single optimized attack point into an MNIST classifier caused the testing error rate to spike from a baseline of 2–5% up to 15–20%, effectively quadrupling error rates. Third, in multi-point poisoning experiments, model degradation scaled steadily as the percentage of contaminated training data increased, consistently undermining system reliability.
These findings have critical operational and security implications. They indicate that machine learning pipelines that automatically retrain on externally ingested data—such as spam reporting feeds or malware collection repositories—face severe vulnerability risks. A small fraction of crafted inputs can critically impair automated classification systems, leading to substantial risks of missed threats, inflated manual review costs, and reduced trust in automated defenses. The article highlights that standard classifiers cannot be assumed secure by default against training-time manipulation.
Organizations utilizing machine learning in security-sensitive workflows should prioritize data integrity verification and develop adversarial defense mechanisms before relying on automated retraining pipelines. Future research must expand from single-point attacks to simultaneous multi-point generation and explore real-world input constraints. Key limitations of this analysis include the assumption of full attacker access to training data distributions and direct control over target labels, which may not hold in human-moderated environments. Nonetheless, confidence in the findings is strong, providing definitive proof that SVM algorithms are fundamentally vulnerable to structured poisoning attacks.
- Paper: Choosing Multiple Parameters for Support Vector Machines, OLIVIER CHAPELLE et al. (2002). This work establishes gradient-based optimization techniques for tuning support vector machine parameters through implicit differentiation of the optimal solution, directly underpinning the poisoning attack's gradient ascent derivation.
- Paper: Support-vector networks, Corinna Cortes et al. (1995). This seminal paper defines the formulation and optimization problem of support vector machines, which serves as the direct target system analyzed for training-data vulnerability.
- Paper: A training algorithm for optimal margin classifiers, B. Boser et al. (1992). This paper establishes the foundational dual optimization and margin-maximization framework of kernel-based support vector machines that the poisoning formulation explicitly exploits.
- Paper: Stability and Generalization, Olivier Bousquet et al. (2002). This paper formalizes algorithmic stability and quantifies how modifying individual training examples alters learned decision functions, providing essential theoretical context for data-manipulation sensitivity.
- Paper: Evasion Attacks against Machine Learning at Test Time, Battista Biggio et al. (2013). Following the source's investigation into training-time poisoning, this complementary work by the same primary authors formulates gradient-based optimization for test-time evasion attacks against machine learning models.
- Paper: BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain, Tianyu Gu et al. (2017). This work extends the paradigm of training-data manipulation by introducing targeted backdoor attacks into deep neural networks during training.
- Paper: Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning, Xinyun Chen et al. (2017). This paper advances data poisoning research by developing black-box targeted backdoor attacks specifically tailored for deep learning architectures.
- Paper: Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples, Nicolas Papernot et al. (2016). This research broadens adversarial vulnerability studies from white-box gradient methods to black-box transfer attacks across diverse classifiers, including support vector machines.
- Paper: The Limitations of Deep Learning in Adversarial Settings, Nicolas Papernot et al. (2015). This paper generalizes the adversarial security threat framework established in shallow models by formalizing systematic adversarial perturbation generation for deep learning architectures.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This text expands the mathematical understanding of gradient-based adversarial vulnerabilities across high-dimensional classification models and proposes adversarial training defenses.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, A. Ma̧dry et al. (2017). This paper advances adversarial robustness by formalizing the defense against gradient-based optimization attacks as a robust saddle-point optimization problem.
