MagNet: A Two-Pronged Defense against Adversarial Examples

Dongyu MengHao Chen

article2017Conference on Computer and Communications Security1,332 citations

Proposes MagNet, an attack-agnostic defense framework that combines detector and reformer networks with randomized diversity to identify and neutralize adversarial inputs without modifying the underlying neural network classifier.

Listen

Deep learning models are increasingly deployed in safety- and security-critical systems such as autonomous vehicles, medical diagnostics, and financial infrastructure. However, these systems are highly vulnerable to adversarial examples—deliberately crafted, tiny perturbations that mislead neural networks into making incorrect classifications while remaining imperceptible to human observers. Most existing defenses either require prior knowledge of specific attack methods, perform poorly against advanced threats, or require retraining the primary classification model, which adds operational complexity and yields limited generalization.

The article introduces and evaluates MagNet, a defensive framework designed to protect neural network classifiers against adversarial examples without altering the underlying classification models or requiring training on attack data. The core objective is to demonstrate that an external, modular defense can generalize across diverse attack strategies, preserve accuracy on legitimate inputs, and effectively mitigate blackbox and graybox threats.

To achieve this, the researchers developed a two-pronged mechanism utilizing separate autoencoder networks trained solely on normal data to approximate the natural data manifold. The architecture deploys detector networks that measure reconstruction error and probability divergence to reject inputs located far from the natural data distribution. It also utilizes a reformer network that reconstructs inputs lying close to the manifold boundary, shifting subtly perturbed inputs back into standard distributions before passing them to the target classifier. The defense was evaluated across standard image recognition benchmarks (MNIST and CIFAR-10) against four advanced adversarial attack algorithms, spanning various perturbation metrics and confidence levels. Additionally, drawing inspiration from cryptographic key randomization, the authors implemented defense diversity by training collections of distinct autoencoders and selecting one at random during runtime to counter graybox attacks.

The experimental findings show that MagNet provides substantial protection across evaluation scenarios. In blackbox settings, where the attacker has no access to defense parameters, MagNet achieved over 99% classification accuracy on nine of ten evaluated attacks on the MNIST dataset, and sustained accuracy between 75% and 100% on the more complex CIFAR-10 benchmark. Against the highly effective Carlini attack, which reduced undefended model accuracy to near 0%, MagNet restored accuracy to over 99% on MNIST and above 80% across all confidence levels on CIFAR-10. This success stems from the complementary operation of its components: the reformer corrects low-confidence, subtle perturbations, while the detectors identify and reject high-confidence, heavily distorted inputs. Under graybox conditions—where attackers understand the defense structure and training process but not the runtime parameter instance—randomly selecting among diversified autoencoders maintained CIFAR-10 classification accuracy above 80%.

These findings indicate that organizations can significantly enhance the operational security and resilience of artificial intelligence systems without modifying proprietary or legacy classifiers. MagNet establishes that defensive systems do not need to model every emerging attack technique to be effective; learning the structure of valid data provides broad protection. This decouples system protection from model development, reducing maintenance costs, engineering friction, and downtime.

For practical implementation, organizations deploying neural networks in untrusted environments should evaluate modular, model-agnostic defense frameworks like MagNet. Security teams should deploy both reconstruction-error and probability-divergence detectors alongside reformer networks to eliminate security gaps across varying perturbation strengths. To mitigate risks where attackers understand the defensive framework, organizations should incorporate runtime diversity by maintaining pools of distinct defensive models.

Decision-makers should note certain limitations and caveats. MagNet does not offer absolute security against complete whitebox attacks, where adversaries possess full knowledge of both classifier and defense parameters, as perfect classifiers off the natural data manifold do not yet exist. Performance is also bounded by the quality of the baseline classifier; lower baseline model accuracy, as observed on CIFAR-10, slightly narrows defensive margins. Continued validation against evolving attack formulations and testing on larger, domain-specific enterprise datasets remain necessary next steps.

Cover for MagNet: A Two-Pronged Defense against Adversarial Examples

Abstract

Deep learning has shown promising results on hard perceptual problems in recent years. However, deep learning systems are found to be vulnerable to small adversarial perturbations that are nearly imperceptible to human. Such specially crafted perturbations cause deep learning systems to output incorrect decisions, with potentially disastrous consequences. These vulnerabilities hinder the deployment of deep learning systems where safety or security is important. Attempts to secure deep learning systems either target specific attacks or have been shown to be ineffective.

In this paper, we propose MagNet, a framework for defending neural network classifiers against adversarial examples. MagNet does not modify the protected classifier or know the process for generating adversarial examples. MagNet includes one or more separate detector networks and a reformer network. Different from previous work, MagNet learns to differentiate between normal and adversarial examples by approximating the manifold of normal examples. Since it does not rely on any process for generating adversarial examples, it has substantial generalization power. Moreover, MagNet reconstructs adversarial examples by moving them towards the manifold, which is effective for helping classify adversarial examples with small perturbation correctly. We discuss the intrinsic difficulty in defending against whitebox attack and propose a mechanism to defend against graybox attack. Inspired by the use of randomness in cryptography, we propose to use diversity to strengthen MagNet. We show empirically that MagNet is effective against most advanced state-of-the-art attacks in blackbox and graybox scenarios while keeping false positive rate on normal examples very low.

Table of Contents

  • 1 Introduction
  • 1.1 Adversarial examples
  • 1.2 Causes of mis-classification and solutions
  • 1.3 Contributions
  • 2 Background and related work
  • 2.1 Deep learning systems in adversarial environments
  • 2.2 Distance metrics
  • 2.3 Existing attacks
  • 2.3.1 Fast gradient sign method(FGSM)
  • 2.3.2 Iterative gradient sign Method
  • 2.3.3 DeepFool
  • 2.3.4 Carlini attack
  • 2.4 Existing defense
  • 2.4.1 Adversarial training
  • 2.4.2 Defensive distillation
  • 2.4.3 Detecting adversarial examples
  • 3 Problem definition
  • 3.1 Adversarial examples
  • 3.2 Defense and evaluation
  • 3.3 Threat model
  • 4 Design
  • 4.1 Detector
  • 4.1.1 Detector based on reconstruction error
  • 4.1.2 Detector based on probability divergence
  • 4.2 Reformer
  • 4.2.1 Noise-based reformer
  • 4.2.2 Autoencoder-based reformer
  • 4.3 Use diversity to mitigate graybox attacks
  • 5 Implementation and Evaluation
  • 5.1 Setup
  • 5.2 Overall performance against blackbox attacks
  • 5.2.1 MNIST
  • 5.2.2 CIFAR-10
  • 5.3 Case study on Carlini attack, why does ?
  • 5.4 Defend against graybox attacks
  • 6 Discussion
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — MagNet Defense Framework

    model/method

    MagNet is a modular defense framework designed to protect neural network classifiers against adversarial examples without modifying the target classifier and without requiring adversarial examples during training.

    Let SS denote the input sample space, and let Nt⊂SN_t \subset S represent the low-dimensional manifold of natural/normal examples for a classification task tt. A target classifier ft:S→Ctf_t: S \to C_t maps an input to a set of classes CtC_t. MagNet treats ftf_t as a black box: it does not alter the classifier's parameters or internal layer representations.

    MagNet operates via two complementary components:

    1. Detector Networks: One or more detectors estimate how far an input candidate x∈Sx \in S lies from the normal data manifold NtN_t. If the estimated distance or output divergence exceeds a pre-set threshold, xx is flagged as adversarial and rejected.

    2. Reformer Network: If an input xx is not rejected by any detector, a reformer function r:S→Ntr: S \to N_t reconstructs xx into an on-manifold or near-manifold representation x′x' that closely approximates xx. The reconstructed input x′x' is then provided to the target classifier ftf_t, yielding the predicted class label ft(x′)f_t(x').

    Because both detectors and reformers are trained exclusively on clean, unperturbed examples to approximate the natural data manifold, the defense is independent of the specific attack mechanism used to craft adversarial examples.

  2. Knowl 2 — Formal Definitions of Adversarial Examples and Defense Correctness

    definition

    Let SS denote the set of all possible examples in the sample space, and let CtC_t be the set of mutually exclusive class labels for a classification task tt. Let Nt={x∈S∣x occurs naturally with regard to task t}N_t = \{x \in S \mid x \text{ occurs naturally with regard to task } t\} denote the subset of normal examples generated with non-negligible probability by the underlying physical process of task tt, assumed to form a low-dimensional manifold in SS.

    • Target Classifier: A function ft:S→Ctf_t: S \to C_t.

    • Ground-Truth Classifier: A function gt:S→Ct∪{⊥}g_t: S \to C_t \cup \{\bot\} representing prevailing human judgment, where ⊥\bot denotes that the input does not belong to the natural data distribution of task tt.

    • Adversarial Example: An example x∈Sx \in S is adversarial for task tt and classifier ftf_t if and only if: ft(x)≠gt(x)andx∈S∖Ntf_t(x) \neq g_t(x) \quad \text{and} \quad x \in S \setminus N_t

    • Defense Function: A defense dft:S→Ct∪{⊥}df_t: S \to C_t \cup \{\bot\} extends the classifier ftf_t to increase robustness against adversarial inputs.

    • Correct Decision Criteria: A defense dftdf_t makes a correct decision on an input xx if either:

      1. x∈Ntx \in N_t (normal example) and dft(x)=gt(x)df_t(x) = g_t(x); or
      2. x∈S∖Ntx \in S \setminus N_t (adversarial example) and either dft(x)=⊥df_t(x) = \bot (rejection) or dft(x)=gt(x)df_t(x) = g_t(x) (correct classification).
  3. Knowl 3 — Threat Models for Adversarial Robustness

    definition

    In adversarial settings, the attacker is assumed to possess complete knowledge of the target classifier ftf_t, including its architecture, parameters, and training procedure. Depending on the attacker's knowledge of the defense mechanism dftdf_t, three threat models are defined:

    • Blackbox Attack: The attacker does not know the parameters or specific instantiation of the defense dftdf_t.

    • Whitebox Attack: The attacker knows all parameters, architectures, and components of the defense dftdf_t, allowing the attacker to treat the composite system dft∘ftdf_t \circ f_t as a single unified differentiable function and generate adversarial examples directly against it.

    • Graybox Attack: The attacker knows the defense architecture, hyper-parameters, training dataset, and training procedure of dftdf_t, but does not know the exact runtime parameter instance (e.g., which specific model from an ensemble of diverse models is chosen randomly at inference time).

  4. Knowl 4 — Autoencoder-Based Reconstruction Error Detector

    model/method

    The reconstruction error detector identifies adversarial examples by measuring the distance between a test input and the learned manifold of normal training data using an autoencoder.

    An autoencoder ae=d∘eae = d \circ e consists of an encoder e:S→He: S \to H mapping input space SS to hidden representation space HH, and a decoder d:H→Sd: H \to S. It is trained on clean training examples Xtrain⊂NtX_{\text{train}} \subset N_t by minimizing mean squared error:

    L(Xtrain)=1∣Xtrain∣∑x∈Xtrain∥x−ae(x)∥22\mathcal{L}(X_{\text{train}}) = \frac{1}{|X_{\text{train}}|} \sum_{x \in X_{\text{train}}} \|x - ae(x)\|_2^2

    Given a test input xx, the reconstruction error is measured using an LpL_p norm:

    E(x)=∥x−ae(x)∥pE(x) = \|x - ae(x)\|_p

    To capture both localized peak perturbations and distributed perturbations, L1L_1 and L2L_2 norms are utilized. An input xx is rejected if E(x)>treE(x) > t_{\text{re}}, where the threshold tret_{\text{re}} is determined using a clean validation set: tret_{\text{re}} is set to the maximum value such that the false positive rate on normal validation examples does not exceed a chosen upper bound tfpt_{\text{fp}}.

  5. Knowl 5 — Temperature-Scaled Probability Divergence Detector

    equation

    The probability divergence detector identifies adversarial examples whose reconstruction errors are small but whose classification predictions diverge significantly before and after autoencoding.

    Let f(x)f(x) be the probability distribution output by applying a temperature-scaled softmax to the classifier logits l(x)∈Rnl(x) \in \mathbb{R}^n for an input xx:

    softmax(l)i=exp⁡(li/T)∑j=1nexp⁡(lj/T)\text{softmax}(l)_i = \frac{\exp(l_i / T)}{\sum_{j=1}^n \exp(l_j / T)}

    where T>1T > 1 is a temperature hyperparameter used to prevent saturation of the softmax function when the top logit dominates.

    Let P=f(x)P = f(x) and Q=f(ae(x))Q = f(ae(x)), where ae(x)ae(x) is the reconstructed output from an autoencoder trained on normal data. The similarity between the two distributions is measured by the Jensen-Shannon divergence:

    JSD(P∥Q)=12DKL(P∥M)+12DKL(Q∥M)\text{JSD}(P \parallel Q) = \frac{1}{2} D_{\text{KL}}(P \parallel M) + \frac{1}{2} D_{\text{KL}}(Q \parallel M)

    where M=12(P+Q)M = \frac{1}{2}(P + Q) and DKL(P∥M)=∑iP(i)log⁡P(i)M(i)D_{\text{KL}}(P \parallel M) = \sum_i P(i) \log \frac{P(i)}{M(i)} is the Kullback-Leibler divergence.

    If JSD(f(x)∥f(ae(x)))\text{JSD}(f(x) \parallel f(ae(x))) exceeds a threshold tjsdt_{\text{jsd}} (calibrated on clean validation data to satisfy a target false positive rate tfpt_{\text{fp}}), xx is rejected as an adversarial example.

  6. Knowl 6 — Autoencoder-Based Reformer

    model/method

    The reformer is a function r:S→Ntr: S \to N_t that projects an untrusted input xx onto or near the manifold of normal examples before the input reaches the target classifier.

    Unlike a naive noise-based reformer r(x)=clip(x+ϵ⋅y)r(x) = \text{clip}(x + \epsilon \cdot y) with y∼N(0,I)y \sim \mathcal{N}(0, I)—which perturbs inputs blindly without learning data geometry—the autoencoder reformer is trained to minimize reconstruction loss on normal examples.

    When evaluated on an adversarial input xx that lies off but close to the manifold boundary (small perturbation), the autoencoder ae(x)ae(x) projects xx back to a nearby point on the normal data manifold. The target classifier receives r(x)=ae(x)r(x) = ae(x) instead of xx. As a result, adversarial perturbations that escape detection are neutralized, restoring the correct classification label while preserving high accuracy on clean inputs.

  7. Knowl 7 — Diversity Regularization Loss for Autoencoders

    equation

    To defend against graybox attacks where the attacker knows the training pipeline and architecture, an ensemble of nn autoencoders is trained simultaneously with a diversity penalty added to the loss function:

    L(x)=∑i=1nMSE(x,aei(x))−α∑i=1nMSE(aei(x),1n∑j=1naej(x))\mathcal{L}(x) = \sum_{i=1}^n \text{MSE}(x, ae_i(x)) - \alpha \sum_{i=1}^n \text{MSE}\left(ae_i(x), \frac{1}{n} \sum_{j=1}^n ae_j(x)\right)

    where:

    • xx is a training sample,
    • aei(x)ae_i(x) is the reconstruction output of the ii-th autoencoder (i∈{1,…,n}i \in \{1, \dots, n\}),
    • MSE(a,b)=∥a−b∥22\text{MSE}(a, b) = \|a - b\|_2^2 is the mean squared error,
    • α>0\alpha > 0 is a regularization hyperparameter governing the tradeoff between individual autoencoder reconstruction fidelity and the mutual diversity (variance across ensemble outputs).

    At test time, the defender randomly samples one autoencoder from {ae1,…,aen}\{ae_1, \dots, ae_n\} for each test input or session. An attacker unaware of the specific random choice must craft perturbations that transfer across all diverse autoencoders simultaneously.

  8. Knowl 8 — Classification Accuracy of MagNet Against Blackbox Attacks

    data/table

    The table compares classification accuracy on adversarial examples with and without MagNet across MNIST and CIFAR-10 datasets under untargeted blackbox attacks. MagNet was not trained on any adversarial examples.

    (a) MNIST
    Attack Norm Parameter No Defense With Defense
    FGSM L∞L^\infty ϵ=0.005\epsilon = 0.005 96.8% 100.0%
    FGSM L∞L^\infty ϵ=0.010\epsilon = 0.010 91.1% 100.0%
    Iterative L∞L^\infty ϵ=0.005\epsilon = 0.005 95.2% 100.0%
    Iterative L∞L^\infty ϵ=0.010\epsilon = 0.010 72.0% 100.0%
    Iterative L2L^2 ϵ=0.5\epsilon = 0.5 86.7% 99.2%
    Iterative L2L^2 ϵ=1.0\epsilon = 1.0 76.6% 100.0%
    DeepFool L∞L^\infty - 19.1% 99.4%
    Carlini L2L^2 - 0.0% 99.5%
    Carlini L∞L^\infty - 0.0% 99.8%
    Carlini L0L^0 - 0.0% 92.0%
    (b) CIFAR-10
    Attack Norm Parameter No Defense With Defense
    FGSM L∞L^\infty ϵ=0.025\epsilon = 0.025 46.0% 99.9%
    FGSM L∞L^\infty ϵ=0.050\epsilon = 0.050 40.5% 100.0%
    Iterative L∞L^\infty ϵ=0.010\epsilon = 0.010 28.6% 96.0%
    Iterative L∞L^\infty ϵ=0.025\epsilon = 0.025 11.1% 99.9%
    Iterative L2L^2 ϵ=0.25\epsilon = 0.25 18.4% 76.3%
    Iterative L2L^2 ϵ=0.50\epsilon = 0.50 6.6% 83.3%
    DeepFool L∞L^\infty - 4.5% 93.4%
    Carlini L2L^2 - 0.0% 93.7%
    Carlini L∞L^\infty - 0.0% 83.0%
    Carlini L0L^0 - 0.0% 77.5%

    On clean test examples, classifier accuracy changes from 99.4% to 99.1% on MNIST, and from 90.6% to 86.8% on CIFAR-10 when MagNet is deployed. On adversarial examples generated by 10 different attacks, MagNet achieves ≥92.0%\ge 92.0\% accuracy on MNIST and ≥76.3%\ge 76.3\% accuracy on CIFAR-10.

  9. Knowl 9 — Complementary Dynamics of Reformer and Detector across Attack Confidence

    empirical result

    When evaluating the Carlini L2L_2 attack across varying confidence parameter values κ∈[0,40]\kappa \in [0, 40] on MNIST and κ∈[0,100]\kappa \in [0, 100] on CIFAR-10, the reformer and detectors exhibit complementary defense profiles:

    1. Low Confidence / Small Perturbation: The adversarial sample is located near the normal manifold boundary. Here, the reformer successfully projects the sample back onto the manifold, restoring correct classification, while the detector exhibits lower detection rates due to small reconstruction errors and low divergence.

    2. High Confidence / Large Perturbation: As confidence increases, perturbation magnitude increases, pushing the adversarial sample further from the normal manifold. The reformer's performance declines, but the detector's rejection rate rises towards 100% due to large reconstruction errors and probability divergences.

    3. Combined Effect: Combining both mechanisms maintains overall classification/detection accuracy ≥99%\ge 99\% across all confidence levels on MNIST and ≥80%\ge 80\% across all confidence levels on CIFAR-10. On CIFAR-10, temperature scaling in the probability divergence detector enhances sensitivity: lower temperature (T=10T=10) yields higher detection at low confidence, while higher temperature (T=40T=40) yields higher detection at high confidence.

  10. Knowl 10 — Mitigation of Graybox Attacks Using Diverse Autoencoders

    data/table

    Eight autoencoders (AA through HH) with identical architectures were trained simultaneously on CIFAR-10 using the diversity loss (with α=0.2\alpha = 0.2). Untargeted Carlini L2L_2 adversarial examples generated against each autoencoder (with detectors omitted to allow whitebox attack synthesis) were evaluated against all autoencoders.

    Test AE Attack Target Autoencoder
    A B C D E F G H
    A 0.0% 92.8% 92.5% 93.1% 91.8% 91.8% 92.5% 93.6%
    B 92.1% 0.0% 92.0% 92.5% 91.4% 92.5% 91.3% 92.5%
    C 93.2% 93.8% 0.0% 92.8% 93.3% 94.1% 92.7% 93.6%
    D 92.8% 92.2% 91.3% 0.0% 91.7% 92.8% 91.2% 93.9%
    E 93.3% 94.0% 93.4% 93.2% 0.0% 93.4% 91.0% 92.8%
    F 92.8% 93.1% 93.2% 93.6% 92.2% 0.0% 92.8% 93.8%
    G 92.5% 93.1% 92.0% 92.2% 90.5% 93.5% 0.1% 93.4%
    H 92.3% 92.0% 91.8% 92.6% 91.4% 92.3% 92.4% 0.0%
    Random 81.1% 81.4% 80.8% 81.3% 80.3% 81.3% 80.5% 81.7%

    When the defense uses the same autoencoder that the attack targeted (the diagonal), accuracy drops to ≈0.0%\approx 0.0\% (whitebox failure). However, transferability between different autoencoders is low: accuracy on non-target autoencoders exceeds 90.5%90.5\%. When the defense randomly selects one autoencoder at test time, accuracy remains above 80.3%80.3\% across all attack targets. On clean test data, the individual autoencoder test accuracies range from 88.7%88.7\% to 89.3%89.3\% (random selection achieves 89.0%89.0\%), compared to 90.6%90.6\% on the baseline unhardened classifier.

  11. Knowl 11 — Foundational Assumptions of the MagNet Framework

    assumption

    The theoretical viability and effectiveness of MagNet rest upon two core assumptions regarding high-dimensional data geometry:

    1. Existence of Manifold Distance Detectors: There exist measurable, learnable detector functions that can approximate the distance between an arbitrary input x∈Sx \in S and the underlying manifold NtN_t of natural data instances.

    2. Existence of Manifold Reformers: There exist transformation functions r:S→Ntr: S \to N_t that map an input x∈Sx \in S to an output x′∈Ntx' \in N_t such that x′x' is perceptibly close to xx while lying strictly closer to (or on) the manifold NtN_t than xx does.

Coverage note — None was omitted; all contributed models, definitions, formulas, empirical evaluations, and foundational assumptions from the paper are fully represented.

References

  1. 1.Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al. End to end learning for self-driving cars. arXiv preprint arXiv:1604 .07316, 2016.
  2. 2.Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In IEEE Symposium on Security and Privacy, 2017.
  3. 3.Shreyansh Daftry, J Andrew Bagnell, and Martial Hebert. Learning transferable policies for monocular reactive mav control. arXiv preprint arXiv:1608.00627, 2016.
  4. 4.Chelsea Finn and Sergey Levine. Deep visual foresight for planning robot motion. arXiv preprint arXiv:1610.00696, 2016.
  5. 5.Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. In International Conference on Learning Representations (ICLR), 2015.
  6. 6.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep Learning. MIT Press, 2016.
  7. 7.Kathrin Grosse, Praveen Manoharan, Nicolas Papernot, Michael Backes, and Patrick McDaniel. On the (statistical) detection of adversarial examples. arXiv preprint arXiv:1702.06280, 2017.
  8. 8.Kathrin Grosse, Nicolas Papernot, Praveen Manoharan, Michael Backes, and Patrick McDaniel. Adversarial perturbations against deep neural networks for malware classification. arXiv preprint arXiv:1606.04435, 2016.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
  10. 10.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503 .02531, 2015.
  11. 11.Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdelrahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al. Deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups. IEEE Signal Processing Magazine, 29(6):82–97, 2012.
  12. 12.Wookhyun Jung, Sangwon Kim, and Sangyong Choi. Poster: deep learning for zero-day flash malware detection. In 36th IEEE Symposium on Security and Privacy, 2015.
  13. 13.Gregory Kahn, Adam Villaflor, Vitchyr Pong, Pieter Abbeel, and Sergey Levine. Uncertainty-aware reinforcement learning for collision avoidance. arXiv preprint arXiv:1702.01182, 2017.
  14. 14.Jernej Kos, Ian Fischer, and Dawn Song. Adversarial examples for generative models. arXiv preprint arXiv:1702.06832, 2017.
  15. 15.Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images, 2009.
  16. 16.Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher. Ask me anything: dynamic memory networks for natural language processing. In International Conference on Machine Learning, pages 1378–1387, 2016.
  17. 17.Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. Adversarial examples in the physical world. CoRR, abs/1607.02533, 2016.
  18. 18.Yann LeCun, Corinna Cortes, and Christopher JC Burges. The mnist database of handwritten digits, 1998.
  19. 19.Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. Delving into transferable adversarial examples and black-box attacks. In International Conference on Learning Representations (ICLR), 2017.
  20. 20.Jan Hendrik Metzen, Tim Genewein, Volker Fischer, and Bastian Bischoff. On detecting adversarial perturbations. In International Conference on Learning Representations (ICLR), April 24–26, 2017.
  21. 21.Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi, and Pascal Frossard. Universal adversarial perturbations. arXiv preprint arXiv:1610.08401, 2016.
  22. 22.Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard. Deepfool: a simple and accurate method to fool deep neural networks. CoRR, abs/1511.04599, 2015.
  23. 23.H. Narayanan and S. Mitter. Sample complexity of testing the manifold hypothesis. In NIPS, 2010.
  24. 24.N. Papernot, P. McDaniel, S. Jha, M. Fredrikson, Z. B. Celik, and A. Swami. The limitations of deep learning in adversarial settings. In IEEE European Symposium on Security and Privacy (EuroSP), 2016.
  25. 25.N. Papernot, P. McDaniel, X. Wu, S. Jha, and A. Swami. Distillation as a defense to adversarial perturbations against deep neural networks. In IEEE Symposium on Security and Privacy, 2016.
  26. 26.Nicolas Papernot, Ian Goodfellow, Ryan Sheatsley, Reuben Feinman, and Patrick McDaniel. Cleverhans v1.0.0: an adversarial machine learning library. arXiv preprint arXiv:1610 .00768, 2016.
  27. 27.Nicolas Papernot, Patrick D. McDaniel, Ananthram Swami, and Richard E. Harang. Crafting adversarial input sequences for recurrent neural networks. CoRR, abs/1604.08275, 2016.
  28. 28.Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z Berkay Celik, and Ananthram Swami. Practical blackbox attacks against machine learning. In Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security, pages 506–519, 2017.
  29. 29.Razvan Pascanu, Jack W Stokes, Hermineh Sanossian, Mady Marinescu, and Anil Thomas. Malware classification with recurrent networks. In Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on, pages 1916–1920. IEEE, 2015.
  30. 30.Uri Shaham, Yutaro Yamada, and Sahand Negahban. Understanding adversarial training: increasing local stability of neural nets through robust optimization. arXiv preprint arXiv :1511.05432, 2015.
  31. 31.Dinggang Shen, Guorong Wu, and Heung-Il Suk. Deep learning in medical image analysis. Annual Review of Biomedical Engineering, (0), 2017.
  32. 32.Justin Sirignano, Apaar Sadhwani, and Kay Giesecke. Deep learning for mortgage risk, 2016.
  33. 33.Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller. Striving for simplicity: the all convolutional net. arXiv preprint arXiv:1412.6806, 2014.
  34. 34.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In International Conference on Learning Representations (ICLR), 2014.
  35. 35.Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning, pages 1096–1103, 2008.
  36. 36.Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, and Pierre-Antoine Manzagol. Stacked denoising autoencoders: learning useful representations in a deep network with a local denoising criterion. Journal of Machine Learning Research, 11(Dec):3371–3408, 2010.
  37. 37.Cihang Xie, Jianyu Wang, Zhishuai Zhang, Yuyin Zhou, Lingxi Xie, and Alan Yuille. Adversarial examples for semantic segmentation and object detection. arXiv preprint arXiv:1703.08603, 2017.

Citation

MLA
Meng, D., and H. Chen. “MagNet: A Two-Pronged Defense Against Adversarial Examples”. arXiv, 2017, http://arxiv.org/abs/1705.09064v2.
APA
Meng, D., & Chen, H. (2017). MagNet: a Two-Pronged Defense against Adversarial Examples. arXiv. http://arxiv.org/abs/1705.09064v2
Chicago
Meng, D., and H. Chen. 2017. “MagNet: A Two-Pronged Defense Against Adversarial Examples”. arXiv. http://arxiv.org/abs/1705.09064v2.
Harvard
Meng, D. and Chen, H. (2017) “MagNet: a Two-Pronged Defense against Adversarial Examples”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1705.09064v2.
Vancouver
1. Meng D, Chen H (2017) MagNet: a Two-Pronged Defense against Adversarial Examples. arXiv

BibTeX

@article{meng2017magnet,
  title = {MagNet: a Two-Pronged Defense against Adversarial Examples},
  author = {Meng, Dongyu and Chen, Hao},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1705.09064v2},
  eprint = {1705.09064}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF