Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples

Nicolas PapernotPatrick McDanielIan Goodfellow

article2016arXiv1,925 citations

Demonstrates how adversarial samples transfer across disparate model architectures to execute practical black-box attacks that successfully deceive commercial machine learning services from Google and Amazon with minimal queries.

Listen

Machine learning systems are increasingly deployed in security-critical environments, including fraud detection, malware filtering, and autonomous navigation. However, these models remain highly vulnerable to adversarial samples—carefully modified inputs designed to trigger incorrect predictions while appearing unaltered to human observers. The article evaluates whether an attacker can systematically execute black-box attacks against remote, proprietary commercial classifiers without knowing their internal architectures, parameters, or training datasets, relying entirely on the transferability of adversarial samples.

To demonstrate this vulnerability, the authors conducted empirical experiments using the MNIST handwritten digit recognition dataset across a broad range of model architectures: deep neural networks, logistic regression, support vector machines, decision trees, nearest neighbors, and multi-model ensembles. They developed new methods to craft adversarial inputs for non-differentiable models such as decision trees and support vector machines. Furthermore, they refined a substitute-model training process using a periodic step size and reservoir sampling, which allows a locally trained substitute to learn the decision boundaries of a remote target classifier with high query efficiency. Finally, they validated the practical threat by launching black-box evasion attacks against production machine learning services hosted by Amazon and Google.

First, the article establishes that adversarial transferability is a widespread phenomenon. Intra-technique transferability was consistently observed across all evaluated algorithms; for example, adversarial samples crafted for one logistic regression model transferred to other logistic regression models at rates exceeding 94%. Second, cross-technique transferability proved equally potent: samples crafted on logistic regression caused support vector machines and decision trees to misclassify 91.43% and 87.42% of inputs, respectively, while multi-model ensembles suffered misclassification rates up to 44.14%. Third, the algorithmic refinements proved highly effective: deep neural networks and logistic regression successfully served as substitute models for almost all target algorithms, matching target predictions on 77% to 89% of test inputs while cutting required queries from over 100,000 to just 2,000 to 3,600. Fourth, practical black-box attacks achieved devastating success against commercial cloud platforms, forcing Amazon Machine Learning to misclassify 96.19% of inputs and Google Cloud Prediction API to misclassify 88.94% of inputs using as few as 800 oracle queries.

These findings demonstrate that standard machine learning deployments present a severe, systemic security risk. Attackers do not need access to model parameters, algorithms, or private training data to reliably deceive production systems, significantly lowering the technical barrier and cost required to compromise operational workflows. Conventional assumptions that black-box deployment or model ensembling provides security through obscurity are invalid. Furthermore, standard defensive countermeasures evaluated in the article, such as retraining models on adversarial inputs, failed to stop attacks on the Google cloud classifier, where misclassification rates remained at 94.2% to 100%.

Organizations deploying machine learning in production should implement rigorous input validation mechanisms analogous to standard security filtering practices like SQL injection defense. Model operators must also monitor query interfaces for suspicious or synthetic querying patterns that indicate an adversary is probing the system to train a substitute model. Cloud service providers should evaluate more robust defenses, such as defensive distillation, and offer visibility into the structural resilience of their automated machine learning pipelines.

The findings provide high confidence regarding the feasibility of black-box transfer attacks across popular model types. However, limitations remain: the empirical evaluation focused primarily on standard image classification benchmarks, and specific internal configurations of proprietary cloud platforms like Google Cloud Prediction remain undisclosed. Decision makers should recognize that developing universal defenses against adversarial inputs remains an open research challenge requiring continuous monitoring and layered operational security.

arXiv: 1605.07277
Cover for Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples

Abstract

Many machine learning models are vulnerable to adversarial examples: inputs that are specially crafted to cause a machine learning model to produce an incorrect output. Adversarial examples that affect one model often affect another model, even if the two models have different architectures or were trained on different training sets, so long as both models were trained to perform the same task. An attacker may therefore train their own substitute model, craft adversarial examples against the substitute, and transfer them to a victim model, with very little information about the victim. Recent work has further developed a technique that uses the victim model as an oracle to label a synthetic training set for the substitute, so the attacker need not even collect a training set to mount the attack. We extend these recent techniques using reservoir sampling to greatly enhance the efficiency of the training procedure for the substitute model. We introduce new transferability attacks between previously unexplored (substitute, victim) pairs of machine learning model classes, most notably SVMs and decision trees. We demonstrate our attacks on two commercial machine learning classification systems from Amazon (96.19% misclassification rate) and Google (88.94%) using only 800 queries of the victim model, thereby showing that existing machine learning approaches are in general vulnerable to systematic black-box attacks regardless of their structure.

Table of Contents

  • 1 Introduction
  • 2 Approach Overview
  • 3 Transferability of Adversarial Samples in Machine Learning
  • 3.1 Experimental Setup
  • 3.2 Intra-technique Transferability
  • 3.3 Cross-technique Transferability
  • 4 Learning Classifier Substitutes by Knowledge Transfer
  • 4.1 Dataset Augmentation for Substitutes
  • 4.2 Deep Neural Network Substitutes
  • 4.3 Logistic Regression Substitutes
  • 4.4 Support Vector Machines Substitutes
  • 5 Black-Box Attacks of Remote Machine Learning Classifiers
  • 5.1 The Oracle Attack Method
  • 5.2 Amazon Web Services Oracle
  • 5.3 Google Cloud Prediction Oracle
  • 6 Adversarial Sample Crafting
  • 6.1 Deep Neural Networks
  • 6.2 Multi-class Logistic Regression
  • 6.3 Nearest Neighbors
  • 6.4 Multi-class Support Vector Machines
  • 6.5 Decision Trees
  • 7 Discussion and Related Work
  • 8 Conclusions
  • References
  • 9 Acknowledgments

Knowls

  1. Knowl 1 — Black-Box Oracle Attack Framework via Substitute Model Training

    model/method

    The black-box oracle attack enables an adversary to generate transferable adversarial samples against a remote target classifier without knowledge of its architecture, parameters, or training data. The adversary is assumed to have only oracle access, allowing them to submit arbitrary inputs x⃗\vec{x} and receive predicted discrete class labels O~(x⃗)\tilde{O}(\vec{x}).

    The attack proceeds in two main phases:

    1. Substitute Model Training: The adversary collects a small initial synthetic/representative dataset S0S_0 (e.g., 100 test samples) and queries the oracle O~\tilde{O} to obtain labels. A local differentiable substitute model ff (such as a Deep Neural Network or Logistic Regression model) is trained on S0S_0. Over ρ\rho iterative augmentation cycles, new synthetic training inputs are generated by evaluating perturbations in the direction of the substitute's Jacobian matrix JfJ_f evaluated at the oracle's predicted label, labeled by querying O~\tilde{O}, and aggregated into Sρ+1S_{\rho+1}.
    2. Adversarial Sample Crafting and Transfer: Once the substitute ff approximates the decision boundaries of O~\tilde{O}, the adversary uses white-box attack methods (such as the Fast Gradient Sign Method) on ff to craft adversarial inputs x⃗∗=x⃗+δx⃗\vec{x}^* = \vec{x} + \delta_{\vec{x}}. By virtue of intra-technique or cross-technique adversarial sample transferability, perturbations misleading ff transfer to mislead the remote oracle O~\tilde{O} with high probability.
  2. Knowl 2 — Jacobian-Based Dataset Augmentation with Reservoir Sampling

    algorithm

    In synthetic substitute model training, standard Jacobian-based augmentation doubles the training dataset size at each iteration ρ\rho, leading to an exponential query complexity n⋅2ρn \cdot 2^\rho for an initial dataset of size nn. To prevent exponential query growth and evade query quotas or detection, reservoir sampling randomly selects a fixed budget of κ\kappa samples from the previous dataset Sρ−1S_{\rho-1} at iterations ρ>σ\rho > \sigma (where the first σ\sigma iterations undergo full augmentation). This reduces the total number of oracle queries to n⋅2σ+κ⋅(ρ−σ)n \cdot 2^\sigma + \kappa \cdot (\rho - \sigma) while preserving the uniform probability 1/∣Sρ−1∣1/|S_{\rho-1}| for any prior input to be augmented.

    Input: Training set Sρ−1S_{\rho-1}, sample budget κ\kappa, substitute Jacobian JfJ_f, step size λρ\lambda_\rho, oracle O~\tilde{O}
    Output: Augmented training set SρS_\rho
    N←∣Sρ−1∣N \leftarrow |S_{\rho-1}|
    Initialize SρS_\rho as array of N+κN + \kappa items
    Sρ[0:N−1]←Sρ−1S_\rho[0 : N - 1] \leftarrow S_{\rho-1}
    for i=0i = 0 to κ−1\kappa - 1 do
        Sρ[N+i]←Sρ−1[i]+λρ⋅sgn(Jf[O~(Sρ−1[i])])S_\rho[N + i] \leftarrow S_{\rho-1}[i] + \lambda_\rho \cdot \text{sgn}(J_f[\tilde{O}(S_{\rho-1}[i])])
    end for
    for i=κi = \kappa to N−1N - 1 do
        r←random integer between 0 and ir \leftarrow \text{random integer between } 0 \text{ and } i
        if r<κr < \kappa then
            Sρ[N+r]←Sρ−1[i]+λρ⋅sgn(Jf[O~(Sρ−1[i])])S_\rho[N + r] \leftarrow S_{\rho-1}[i] + \lambda_\rho \cdot \text{sgn}(J_f[\tilde{O}(S_{\rho-1}[i])])
        end if
    end for
    return SρS_\rho
  3. Knowl 3 — Periodic Step Size for Jacobian-Based Dataset Augmentation

    equation

    To enhance the quality of the decision boundary approximation in substitute model training, the augmentation step size parameter λρ\lambda_\rho alternates periodically between positive and negative directions across augmentation iterations ρ∈N\rho \in \mathbb{N}:

    λρ=λ⋅(−1)⌊ρτ⌋\lambda_\rho = \lambda \cdot (-1)^{\left\lfloor \frac{\rho}{\tau} \right\rfloor}

    where λ>0\lambda > 0 represents the base step amplitude (e.g., λ=0.1\lambda = 0.1), and τ∈N+\tau \in \mathbb{N}^+ denotes the iteration period (e.g., τ=3\tau = 3). Multiplying the step size by −1-1 every τ\tau iterations forces the synthetic points to explore input space on both sides of the local decision boundaries, improving the proportion of matched labels between the substitute model and the target oracle compared to a constant positive step size.

  4. Knowl 4 — Adversarial Sample Transferability Metric and Cross-Technique Dynamics

    definition

    Adversarial sample transferability formalizes the property where an adversarial sample crafted to mislead a source classifier ff also induces an error in a target classifier f′f'. For an input distribution XX and an adversarial perturbation δx⃗\delta_{\vec{x}} computed such that f(x⃗+δx⃗)≠f(x⃗)f(\vec{x} + \delta_{\vec{x}}) \neq f(\vec{x}), the transferability rate is defined as:

    ΩX(f,f′)=∣{x⃗∈X:f′(x⃗)≠f′(x⃗+δx⃗)}∣∣X∣\Omega_X(f, f') = \frac{\left| \{ \vec{x} \in X : f'(\vec{x}) \neq f'(\vec{x} + \delta_{\vec{x}}) \} \right|}{|X|}

    Transferability is partitioned into two variants:

    1. Intra-technique transferability: ff and f′f' are trained using the same machine learning algorithm (e.g., both are DNNs or both are Decision Trees) but with different random initializations or disjoint training subsets.
    2. Cross-technique transferability: ff and f′f' are trained using fundamentally different machine learning algorithms (e.g., ff is a DNN and f′f' is an SVM or Decision Tree).

    Empirical evaluations show that differentiable models (DNN, Logistic Regression) exhibit high intra-technique transferability (>49% for DNN, >94% for LR). Cross-technique transferability is substantial across diverse model families: adversarial samples crafted on Logistic Regression transfer to SVMs (91.43%), Decision Trees (87.42%), and multi-model Ensembles (44.14%). Decision Trees are the most vulnerable target to cross-technique attacks (misclassification rates of 47.20%–89.29%), while DNNs are the most resilient target (0.82%–38.27%).

  5. Knowl 5 — Black-Box Attacks on Amazon and Google Cloud Machine Learning Platforms

    empirical result

    Black-box adversarial attacks were conducted against remote commercial machine learning classifiers hosted on Amazon Web Services (Amazon Machine Learning) and Google Cloud Prediction API, trained on the MNIST handwritten digit dataset. Substitutes were initialized with only 100 test samples and trained via Jacobian-based augmentation with Periodic Step Size (PSS, τ=3\tau=3) and Reservoir Sampling (RS, σ=3,κ=400\sigma=3, \kappa=400). Adversarial samples were generated using the Fast Gradient Sign Method with perturbation magnitude ε=0.3\varepsilon = 0.3.

    Target Oracle Substitute Type Iterations Query Count Misclassification Rate
    Amazon ML Logistic Regression (LR) ρ=3\rho = 3 800 96.19%
    Amazon ML DNN ρ=3\rho = 3 800 87.44%
    Amazon ML DNN ρ=6\rho = 6 6,400 96.78%
    Amazon ML DNN (PSS + RS) ρ=6\rho = 6 2,000 95.68%
    Amazon ML LR (PSS + RS) ρ=6\rho = 6 2,000 95.83%
    Google Cloud Logistic Regression (LR) ρ=3\rho = 3 800 88.94%
    Google Cloud DNN ρ=3\rho = 3 800 84.50%
    Google Cloud DNN ρ=6\rho = 6 6,400 97.17%
    Google Cloud DNN (PSS + RS) ρ=6\rho = 6 2,000 91.57%
    Google Cloud LR (PSS + RS) ρ=6\rho = 6 2,000 97.72%

    The target classifier on Amazon achieved 92.17% baseline accuracy (employing multinomial logistic regression), and Google Cloud achieved 92.00% baseline accuracy. Using only 800 black-box queries, LR substitutes forced 96.19% misclassification on Amazon and 88.94% on Google. Combining PSS and RS reduced the queries required at ρ=6\rho=6 by more than a factor of 3 (from 6,400 to 2,000) while maintaining misclassification rates above 91%–97%.

  6. Knowl 6 — Adversarial Sample Crafting for Linear Multiclass Support Vector Machines

    model/method

    For a multiclass linear Support Vector Machine (SVM) composed of binary classifiers fk(x⃗)=sgn(w⃗[k]⋅x⃗+bk)f_k(\vec{x}) = \text{sgn}(\vec{w}[k] \cdot \vec{x} + b_k) trained under a one-vs-the-rest scheme, an adversarial perturbation can be computed in closed form without iterative optimization.

    Given a legitimate sample x⃗\vec{x} assigned class k=arg⁡max⁡j(w⃗[j]⋅x⃗+bj)k = \arg\max_j (\vec{w}[j] \cdot \vec{x} + b_j), the adversarial sample x⃗∗\vec{x}^* is crafted by displacing x⃗\vec{x} in the direction orthogonal to the separating hyperplane of the binary subclassifier fkf_k:

    x⃗∗=x⃗−ε⋅w⃗[k]∥w⃗[k]∥\vec{x}^* = \vec{x} - \varepsilon \cdot \frac{\vec{w}[k]}{\|\vec{w}[k]\|}

    where w⃗[k]\vec{w}[k] is the weight vector of the binary SVM for class kk, ∥⋅∥\|\cdot\| is the Frobenius/Euclidean norm, and ε>0\varepsilon > 0 is the input perturbation parameter controlling the distortion magnitude.

  7. Knowl 7 — Adversarial Sample Crafting for Decision Trees

    algorithm

    Decision trees partition the input space via axis-aligned orthogonal splits and are non-differentiable. To craft adversarial samples against a decision tree TT, an adversary exploits its hierarchical structure by searching for an alternate target leaf yielding a different class in the immediate neighborhood of the legitimate sample's leaf, and then minimally modifying the sample's features to satisfy the decision path leading to that target leaf.

    Input: Decision tree TT, sample x⃗\vec{x}, correct class legitimate_class\text{legitimate\_class}
    Output: Adversarial sample x⃗∗\vec{x}^*
    x⃗∗←x⃗\vec{x}^* \leftarrow \vec{x}
    legit_leaf←find leaf in T corresponding to x⃗\text{legit\_leaf} \leftarrow \text{find leaf in } T \text{ corresponding to } \vec{x}
    ancestor←legit_leaf\text{ancestor} \leftarrow \text{legit\_leaf}
    components←[]\text{components} \leftarrow []
    while predict(T,x⃗∗)==legitimate_class\text{predict}(T, \vec{x}^*) == \text{legitimate\_class} do
        if ancestor==ancestor.parent.left\text{ancestor} == \text{ancestor.parent.left} then
            advers_leaf←find leaf under ancestor.parent.right\text{advers\_leaf} \leftarrow \text{find leaf under ancestor.parent.right}
        else
            advers_leaf←find leaf under ancestor.parent.left\text{advers\_leaf} \leftarrow \text{find leaf under ancestor.parent.left}
        end if
        components←nodes from legit_leaf to advers_leaf\text{components} \leftarrow \text{nodes from legit\_leaf to advers\_leaf}
        ancestor←ancestor.parent\text{ancestor} \leftarrow \text{ancestor.parent}
    end while
    for each node i∈componentsi \in \text{components} do
        perturb x⃗∗[i]\vec{x}^*[i] to change the condition outcome of node ii
    end for
    return x⃗∗\vec{x}^*
  8. Knowl 8 — Jacobian Formulation for Multi-Class Logistic Regression Substitutes

    equation

    When using a multi-class logistic regression model f:x⃗↦[ew⃗j⋅x⃗∑l=1New⃗l⋅x⃗]j∈{1,…,N}f: \vec{x} \mapsto \left[ \frac{e^{\vec{w}_j \cdot \vec{x}}}{\sum_{l=1}^N e^{\vec{w}_l \cdot \vec{x}}} \right]_{j \in \{1, \dots, N\}} as a substitute model for synthetic dataset augmentation, the (i,j)(i, j)-th component of the Jacobian matrix Jf(x⃗)J_f(\vec{x}) (representing the derivative of class probability output jj with respect to input feature ii) is computed analytically as:

    Jf(x⃗)[i,j]=wj[i]ew⃗j⋅x⃗−∑l=1Nwl[i]⋅ew⃗l⋅x⃗(∑l=1New⃗l⋅x⃗)2J_f(\vec{x})[i, j] = \frac{w_j[i] e^{\vec{w}_j \cdot \vec{x}} - \sum_{l=1}^N w_l[i] \cdot e^{\vec{w}_l \cdot \vec{x}}}{\left( \sum_{l=1}^N e^{\vec{w}_l \cdot \vec{x}} \right)^2}

    where w⃗j\vec{w}_j is the weight vector corresponding to class jj, wj[i]w_j[i] is the weight assigned to input feature ii for class jj, and NN is the total number of classes.

  9. Knowl 9 — Fidelity of Substitute Models Across Diverse Oracle Architectures

    data/table

    The ability of Deep Neural Network (DNN) and Logistic Regression (LR) substitute models to approximate various target oracle architectures (DNN, LR, SVM, Decision Tree [DT], and 1-Nearest Neighbor [kNN]) was evaluated on MNIST test data after ρ=9\rho = 9 augmentation iterations across three configurations: Vanilla, Periodic Step Size (PSS, τ=3\tau=3), and PSS combined with Reservoir Sampling (RS, σ=3,κ=400\sigma=3, \kappa=400).

    Substitute Configuration DNN Oracle LR Oracle SVM Oracle DT Oracle kNN Oracle
    DNN 78.01% 82.17% 79.68% 62.75% 81.83%
    DNN + PSS 89.28% 89.16% 83.79% 61.10% 85.67%
    DNN + PSS + RS 82.90% 83.33% 77.22% 48.62% 82.46%
    LR 64.93% 72.00% 71.56% 38.44% 70.74%
    LR + PSS 69.20% 84.01% 82.19% 34.14% 71.02%
    LR + PSS + RS 67.85% 78.94% 79.20% 41.93% 70.92%

    Key takeaways:

    1. PSS consistently increases label agreement between substitute models and target oracles by up to 11.27 percentage points (e.g., DNN substitute matching DNN oracle rose from 78.01% to 89.28%).
    2. RS reduces query complexity by 96.5% (from 102,400 to 3,600 queries at ρ=9\rho=9) with minimal loss in label matching fidelity.
    3. Both DNN and LR substitutes readily approximate smooth and lazy classifiers (matching >78%–89% of DNN, LR, SVM, and kNN labels), but struggle on non-differentiable Decision Trees (matching only 34%–62%). Conversely, SVM-based substitutes fail to extract knowledge from non-SVM oracles (matching only ~12% of DNN and LR oracle labels).
  10. Knowl 10 — Ineffectiveness of Adversarial Retraining as a Defense for Cloud Machine Learning APIs

    limitation

    Adversarial retraining—augmenting the training dataset with adversarial samples during training—was evaluated as a potential defense for the Google Cloud Prediction API. An MNIST model retrained with adversarial samples added to its training set achieved a legitimate test accuracy of 91.25%.

    However, when a new black-box DNN substitute model was trained against this hardened oracle without knowing its internal defense, the crafted adversarial samples still achieved a misclassification rate of 94.2% after ρ=3\rho = 3 substitute augmentation iterations and 100% after ρ=6\rho = 6 iterations. This demonstrates that standard adversarial data augmentation is ineffective at protecting black-box cloud classifiers against substitute-based transfer attacks, likely due to the limited capacity and shallowness of the underlying commercial models.

Coverage note — None was omitted; all key theoretical formulations, algorithmic innovations, empirical transferability matrices, black-box cloud platform attack experiments, and defense evaluations have been captured.

References

  1. 1.M. Barreno, B. Nelson, R. Sears, A. D. Joseph, and J. D. Tygar. Can machine learning be secure? In Proceedings of the 2006 ACM Symposium on Information, computer and communications security, pages 16–25. ACM, 2006.
  2. 2.E. Battenberg, S. Dieleman, and al. Lasagne: Lightweight library to build and train neural networks in theano, 2015.
  3. 3.J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, and al. Theano: a cpu and gpu math expression compiler. In Proceedings of the Python for scientific computing conference (SciPy), volume 4, page 3. Austin, TX, 2010.
  4. 4.B. Biggio, I. Corona, and al. Evasion attacks against machine learning at test time. In Machine Learning and Knowledge Discovery in Databases, pages 387–402. Springer, 2013.
  5. 5.B. Biggio, G. Fumera, and F. Roli. Security evaluation of pattern classifiers under attack. Knowledge and Data Engineering, IEEE Transactions on, 26(4):984–996, 2014.
  6. 6.B. Biggio, B. Nelson, and P. Laskov. Support vector machines under adversarial label noise. In ACML, pages 97–112, 2011.
  7. 7.B. Biggio, B. Nelson, and L. Pavel. Poisoning attacks against support vector machines. In Proceedings of the 29th International Conference on Machine Learning, 2012.
  8. 8.C. M. Bishop. Pattern recognition. Machine Learning, 2006.
  9. 9.C. Bucila, R. Caruana, and A. Niculescu-Mizil. Model compression. In Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 535–541. ACM, 2006.
  10. 10.T. Chen, I. Goodfellow, and J. Shlens. Net2net: Accelerating learning via knowledge transfer. In Proceedings of the 2016 International Conference on Learning Representations. Computational and Biological Learning Society, 2016.
  11. 11.I. Goodfellow, Y. Bengio, and A. Courville. Deep learning. Book in preparation for MIT Press, 2016.
  12. 12.I. J. Goodfellow, J. Shlens, and C. Szegedy. Explaining and harnessing adversarial examples. In Proceedings of the 2015 International Conference on Learning Representations. Computational and Biological Learning Society, 2015.
  13. 13.G. Hinton, O. Vinyals, and J. Dean. Distilling the knowledge in a neural network. In Deep Learning and Representation Learning Workshop at NIPS 2014. arXiv preprint arXiv:1503.02531, 2014.
  14. 14.L. Huang, A. D. Joseph, B. Nelson, B. I. Rubinstein, and J. Tygar. Adversarial machine learning. In Proceedings of the 4th ACM workshop on Security and artificial intelligence, pages 43–58. ACM, 2011.
  15. 15.M. Kloft and P. Laskov. Online anomaly detection under adversarial impact. In International Conference on Artificial Intelligence and Statistics, pages 405–412, 2010.
  16. 16.Y. LeCun and C. Cortes. The mnist database of handwritten digits, 1998.
  17. 17.P. McDaniel, N. Papernot, and Z. B. Celik. Machine Learning in Adversarial Settings. IEEE Security & Privacy Magazine, 14(3), May/June 2016.
  18. 18.K. P. Murphy. Machine learning: a probabilistic perspective. MIT press, 2012.
  19. 19.N. Papernot, P. McDaniel, and al. The limitations of deep learning in adversarial settings. In Proceedings of the 1st IEEE European Symposium on Security and Privacy. IEEE, 2016.
  20. 20.N. Papernot, P. McDaniel, I. Goodfellow, S. Jha, and al. Practical black-box attacks against deep learning systems using adversarial examples. arXiv preprint arXiv:1602.02697, 2016.
  21. 21.N. Papernot, P. McDaniel, X. Wu, S. Jha, and A. Swami. Distillation as a defense to adversarial perturbations against deep neural networks. In Proceedings of the 37th IEEE Symposium on Security and Privacy. IEEE, 2016.
  22. 22.C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, , et al. Intriguing properties of neural networks. In Proceedings of the 2014 International Conference on Learning Representations. Computational and Biological Learning Society, 2014.
  23. 23.J. S. Vitter. Random sampling with a reservoir. ACM Transactions on Mathematical Software (TOMS), 11(1):37–57, 1985.
  24. 24.D. Warde-Farley and I. Goodfellow. Adversarial perturbations of deep neural networks. In T. Hazan, G. Papandreou, and D. Tarlow, editors, Advanced Structured Prediction. 2016.

Citation

MLA
Papernot, N., et al. “Transferability in Machine Learning: From Phenomena to Black-Box Attacks Using Adversarial Samples”. arXiv, 2016, http://arxiv.org/abs/1605.07277v1.
APA
Papernot, N., McDaniel, P., & Goodfellow, I. (2016). Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples. arXiv. http://arxiv.org/abs/1605.07277v1
Chicago
Papernot, N., P. McDaniel, and I. Goodfellow. 2016. “Transferability in Machine Learning: From Phenomena to Black-Box Attacks Using Adversarial Samples”. arXiv. http://arxiv.org/abs/1605.07277v1.
Harvard
Papernot, N., McDaniel, P. and Goodfellow, I. (2016) “Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1605.07277v1.
Vancouver
1. Papernot N, McDaniel P, Goodfellow I (2016) Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples. arXiv

BibTeX

@article{papernot2016transferability,
  title = {Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples},
  author = {Papernot, Nicolas and McDaniel, Patrick and Goodfellow, Ian},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1605.07277v1},
  eprint = {1605.07277}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission