Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples
Nicolas PapernotPatrick McDanielIan Goodfellow
Demonstrates how adversarial samples transfer across disparate model architectures to execute practical black-box attacks that successfully deceive commercial machine learning services from Google and Amazon with minimal queries.
Machine learning systems are increasingly deployed in security-critical environments, including fraud detection, malware filtering, and autonomous navigation. However, these models remain highly vulnerable to adversarial samples—carefully modified inputs designed to trigger incorrect predictions while appearing unaltered to human observers. The article evaluates whether an attacker can systematically execute black-box attacks against remote, proprietary commercial classifiers without knowing their internal architectures, parameters, or training datasets, relying entirely on the transferability of adversarial samples.
To demonstrate this vulnerability, the authors conducted empirical experiments using the MNIST handwritten digit recognition dataset across a broad range of model architectures: deep neural networks, logistic regression, support vector machines, decision trees, nearest neighbors, and multi-model ensembles. They developed new methods to craft adversarial inputs for non-differentiable models such as decision trees and support vector machines. Furthermore, they refined a substitute-model training process using a periodic step size and reservoir sampling, which allows a locally trained substitute to learn the decision boundaries of a remote target classifier with high query efficiency. Finally, they validated the practical threat by launching black-box evasion attacks against production machine learning services hosted by Amazon and Google.
First, the article establishes that adversarial transferability is a widespread phenomenon. Intra-technique transferability was consistently observed across all evaluated algorithms; for example, adversarial samples crafted for one logistic regression model transferred to other logistic regression models at rates exceeding 94%. Second, cross-technique transferability proved equally potent: samples crafted on logistic regression caused support vector machines and decision trees to misclassify 91.43% and 87.42% of inputs, respectively, while multi-model ensembles suffered misclassification rates up to 44.14%. Third, the algorithmic refinements proved highly effective: deep neural networks and logistic regression successfully served as substitute models for almost all target algorithms, matching target predictions on 77% to 89% of test inputs while cutting required queries from over 100,000 to just 2,000 to 3,600. Fourth, practical black-box attacks achieved devastating success against commercial cloud platforms, forcing Amazon Machine Learning to misclassify 96.19% of inputs and Google Cloud Prediction API to misclassify 88.94% of inputs using as few as 800 oracle queries.
These findings demonstrate that standard machine learning deployments present a severe, systemic security risk. Attackers do not need access to model parameters, algorithms, or private training data to reliably deceive production systems, significantly lowering the technical barrier and cost required to compromise operational workflows. Conventional assumptions that black-box deployment or model ensembling provides security through obscurity are invalid. Furthermore, standard defensive countermeasures evaluated in the article, such as retraining models on adversarial inputs, failed to stop attacks on the Google cloud classifier, where misclassification rates remained at 94.2% to 100%.
Organizations deploying machine learning in production should implement rigorous input validation mechanisms analogous to standard security filtering practices like SQL injection defense. Model operators must also monitor query interfaces for suspicious or synthetic querying patterns that indicate an adversary is probing the system to train a substitute model. Cloud service providers should evaluate more robust defenses, such as defensive distillation, and offer visibility into the structural resilience of their automated machine learning pipelines.
The findings provide high confidence regarding the feasibility of black-box transfer attacks across popular model types. However, limitations remain: the empirical evaluation focused primarily on standard image classification benchmarks, and specific internal configurations of proprietary cloud platforms like Google Cloud Prediction remain undisclosed. Decision makers should recognize that developing universal defenses against adversarial inputs remains an open research challenge requiring continuous monitoring and layered operational security.
- Paper: The Limitations of Deep Learning in Adversarial Settings, Nicolas Papernot et al. (2015). This paper introduced Jacobian-based saliency map attacks (JSMA) and threat models in deep learning, providing foundational adversarial generation techniques adapted in the source's substitute-model attacks.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This foundational work introduced the Fast Gradient Sign Method and established early observations of adversarial transferability that the source directly builds upon.
- Paper: Intriguing properties of neural networks, Christian Szegedy et al. (2014). This foundational paper first uncovered the vulnerability of neural networks to imperceptible adversarial perturbations and observed cross-model generalization.
- Paper: Evasion Attacks against Machine Learning at Test Time, Battista Biggio et al. (2013). This paper formulated early evasion attacks against machine learning algorithms using surrogate models under limited adversary knowledge.
- Paper: DeepFool: A Simple and Accurate Method to Fool Deep Neural Networks, Seyed-Mohsen Moosavi-Dezfooli et al. (2015). This work developed DeepFool for computing minimal adversarial perturbations, serving as an important comparative attack methodology for generating transferable samples.
- Paper: Distillation as a Defense to Adversarial Perturbations Against Deep Neural Networks, Nicolas Papernot et al. (2015). This paper introduced defensive distillation as an adversarial defense, which became a primary target for black-box transferability evaluation.
- Paper: Practical Black-Box Attacks against Machine Learning, Nicolas Papernot et al. (2017). This subsequent paper formalizes practical black-box attacks using synthetic dataset augmentation and substitute models against cloud-based machine learning APIs.
- Paper: Delving into Transferable Adversarial Examples and Black-box Attacks, Yanpei Liu et al. (2016). This work extends black-box transferability to large-scale ImageNet vision models and commercial platforms using ensemble-based optimization.
- Paper: ZOO: Zeroth Order Optimization Based Black-box Attacks to Deep Neural Networks without Training Substitute Models, Pin-Yu Chen et al. (2017). This paper develops zeroth-order optimization to mount black-box attacks directly without needing to train the substitute models introduced in transferability research.
- Paper: Ensemble Adversarial Training: Attacks and Defenses, Florian Tramèr et al. (2018). This study analyzes the dynamics of transferred adversarial examples and introduces ensemble adversarial training to defend against transferability attacks.
- Paper: Towards Evaluating the Robustness of Neural Networks, Nicholas Carlini et al. (2016). This work establishes standard optimization-based attacks (C&W) and systematically evaluates the cross-model transferability of high-confidence adversarial examples.
- Paper: Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods, Nicholas Carlini et al. (2017). This paper evaluates ten adversarial detection defenses, utilizing transferability-based black-box attacks to demonstrate the vulnerability of adversarial detectors.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, A. Ma̧dry et al. (2017). This paper provides a robust optimization framework using projected gradient descent to train models resistant to first-order and transferred adversarial attacks.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). This work investigates why adversarial examples transfer so effectively across models by demonstrating that they stem from dataset-intrinsic non-robust features.
- Paper: Universal Adversarial Perturbations, Seyed-Mohsen Moosavi-Dezfooli et al. (2016). This paper generalizes sample-specific perturbations by finding universal adversarial perturbations that generalize across entire datasets and different model architectures.
- Paper: Stealing Machine Learning Models via Prediction APIs, Florian Tramèr et al. (2016). This study examines model extraction attacks against prediction APIs, closely complementing the black-box substitute model generation paradigm.
