Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning

Xinyun ChenChang LiuBo LiKimberly LuDawn Song

article2017arXiv2,352 citations

Demonstrates that deep neural networks can be compromised with targeted backdoors by injecting as few as fifty stealthy poisoned samples into the training data without any knowledge of the underlying model or training pipeline.

Listen

Deep learning models are increasingly deployed in high-stakes, security-critical environments such as facial recognition for physical access control, passport verification, and financial authentication. However, training these complex systems requires massive amounts of data, creating serious supply chain and insider risk vulnerabilities. The article evaluates whether an adversary can successfully manipulate a deep learning model by covertly inserting a small amount of poisoned training data, thereby creating a targeted backdoor. The core objective was to demonstrate the feasibility of backdoor data poisoning under a strict black-box threat model, where the attacker has no knowledge of the model architecture or benign training data, injects an imperceptibly small volume of corrupted samples, and uses keys that are stealthy and physically deployable.

To test this threat, the researchers developed two primary attack strategies: input-instance-key attacks, which trigger the backdoor using variations of a specific input image, and pattern-key attacks, which embed specific visual triggers such as blended watermarks or wearable accessories. They evaluated these techniques on two state-of-the-art face recognition architectures—DeepID trained from scratch and a fine-tuned VGG-Face model—using the YouTube Aligned Face dataset comprising approximately 600,000 images across 1,283 identities. The study evaluated attack success rates, baseline model accuracy on normal data, selectivity against incorrect keys, and the viability of physical attacks using real commodity eyeglasses and sunglasses photographed from multiple angles.

The findings establish that deep learning systems are exceptionally vulnerable to targeted backdoor poisoning even under minimal adversary capabilities. First, for input-instance-key attacks, injecting as few as 5 poisoned samples into a training set of hundreds of thousands of images achieved a 100% attack success rate without degrading the baseline test accuracy of roughly 97.5% to 97.8%. Second, for pattern-key attacks, injecting approximately 50 to 115 poisoned samples—representing a tiny fraction of the total dataset—achieved attack success rates exceeding 90% while keeping the backdoor pattern virtually imperceptible to human reviewers. Third, physical real-world attacks proved highly effective: individuals wearing off-the-shelf sunglasses or reading glasses successfully triggered the backdoor across different camera angles. Finally, the attacks proved highly selective, yielding a 0% false trigger rate when presented with incorrect keys, and standard defenses such as label distribution monitoring, statistical outlier detection, and pre-training on clean auxiliary data failed to mitigate the risk.

These results demonstrate that deep learning authentication systems face severe security and compliance risks from insider threats or compromised data pipelines. Because backdoored models maintain standard performance on legitimate inputs and exhibit no obvious statistical anomalies, organizations cannot rely on traditional model validation, test accuracy metrics, or basic data filtering to detect tampering. The ability to execute physical impersonation attacks using regular commercial items significantly increases the operational danger for biometric access control and identity management systems.

Organizations developing or deploying deep learning models in sensitive roles must recognize that existing baseline defenses and outlier detection mechanisms are insufficient. Decision-makers should prioritize securing data supply chains, implementing strict data provenance and access controls for labeling personnel, and investing in research toward robust backdoor detection methods. While the experimental findings strongly demonstrate vulnerabilities in facial recognition frameworks, further evaluation is warranted to measure how these poisoning strategies perform across diverse operational domains, such as audio recognition, natural language processing, and autonomous navigation.

arXiv: 1712.05526
  • Paper: How To Backdoor Federated Learning, Eugene Bagdasaryan et al. (2018). This paper directly extends data poisoning and backdoor concepts to decentralized federated learning systems.
Cover for Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning

Abstract

Deep learning models have achieved high performance on many tasks, and thus have been applied to many security-critical scenarios. For example, deep learning-based face recognition systems have been used to authenticate users to access many security-sensitive applications like payment apps. Such usages of deep learning systems provide the adversaries with sufficient incentives to perform attacks against these systems for their adversarial purposes. In this work, we consider a new type of attacks, called backdoor attacks, where the attacker's goal is to create a backdoor into a learning-based authentication system, so that he can easily circumvent the system by leveraging the backdoor. Specifically, the adversary aims at creating backdoor instances, so that the victim learning system will be misled to classify the backdoor instances as a target label specified by the adversary. In particular, we study backdoor poisoning attacks, which achieve backdoor attacks using poisoning strategies. Different from all existing work, our studied poisoning strategies can apply under a very weak threat model: (1) the adversary has no knowledge of the model and the training set used by the victim system; (2) the attacker is allowed to inject only a small amount of poisoning samples; (3) the backdoor key is hard to notice even by human beings to achieve stealthiness. We conduct evaluation to demonstrate that a backdoor adversary can inject only around 50 poisoning samples, while achieving an attack success rate of above 90%. We are also the first work to show that a data poisoning attack can create physically implementable backdoors without touching the training process. Our work demonstrates that backdoor poisoning attacks pose real threats to a learning system, and thus highlights the importance of further investigation and proposing defense strategies against them.

Table of Contents

  • I. INTRODUCTION
  • A. A Motivating Example: Face Recognition Systems
  • B. A Realistic Threat Model and Attack Goals
  • C. Contributions
  • II. BACKDOOR POISONING ATTACKS
  • A. Backdoor Attack in a Learning System
  • B. Backdoor Adversary Using Data Poisoning
  • III. BACKDOOR POISONING ATTACK STRATEGIES
  • A. Input-instance-key strategies
  • B. Pattern-key strategies
  • IV. EVALUATION SETUP
  • A. Dataset.
  • B. Models
  • C. Metrics
  • V. EVALUATION OF BACKDOOR POISONING ATTACKS
  • A. Evaluation of the input-instance-key strategy
  • B. Evaluation of the Blended Injection strategy
  • C. Evaluation of the Accessory Injection strategy
  • D. Evaluation of the Blended Accessory Injection strategy
  • VI. EVALUATION OF PHYSICAL ATTACKS
  • VII. EVALUATION OF POTENTIAL (FAILING) DEFENSES
  • A. Detection of label distribution
  • B. Outlier detector-based defense
  • C. Defense with auxiliary pristine data
  • VIII. RELATED WORK
  • IX. CONCLUSION AND FUTURE WORK
  • ACKNOWLEDGMENT
  • REFERENCES
  • APPENDIX
  • A. Examples of the attacks
  • B. Model details
  • C. More results of physical attacks

Knowls

  1. Knowl 1 — Formal Framework and Threat Model for Backdoor Poisoning Attacks

    definition

    A backdoor poisoning attack against a machine learning classification model fθ:X→Yf_\theta: \mathcal{X} \to \mathcal{Y} embeds a targeted backdoor into the model solely by injecting a small quantity nn of poisoned training instances into the training dataset D={(xi,yi)∈X×Y∣i=1,…,N}\mathcal{D} = \{(x_i, y_i) \in \mathcal{X} \times \mathcal{Y} \mid i = 1, \dots, N\}, yielding a poisoned training set Dpoison=D∪{(xip,yip)}i=1n\mathcal{D}_{\text{poison}} = \mathcal{D} \cup \{(x_i^p, y_i^p)\}_{i=1}^n.

    The attack is parameterized by a target label yt∈Yy^t \in \mathcal{Y}, a backdoor key k∈Kk \in \mathcal{K} (where the key space K\mathcal{K} may or may not overlap with the input space X\mathcal{X}), and a backdoor-instance-generation function Σ:K→2X\Sigma: \mathcal{K} \to 2^\mathcal{X} mapping key kk to a subspace of backdoor instances. For all injected poisoning samples, the assigned label is set to the target label (yip=yty_i^p = y^t).

    The adversary operates under a black-box threat model with the following properties:

    1. Black-box access: The adversary has no knowledge of the model architecture ff or its learned parameters θ\theta.
    2. Unawareness of pristine training data: The adversary constructs poisoning samples without access to the pristine training set D\mathcal{D}.
    3. Limited injection volume: n≪Nn \ll N (e.g., n≈5–50n \approx 5\text{--}50 samples when N≈600,000N \approx 600{,}000).
    4. Stealthiness: The backdoor key is visually inconspicuous or imperceptible to human inspection, and the model's standard accuracy on pristine test data T\mathcal{T} is preserved.

    The attack evaluates three performance criteria:

    • Attack Success Rate (ASR): Pr⁡(fθ(xb)=yt)\Pr(f_\theta(x^b) = y^t) for backdoor instances xb∈Σ(k)x^b \in \Sigma(k), where the prediction probability must meet or exceed a confidence acceptance threshold (e.g., 0.850.85).
    • Standard Test Accuracy: Accuracy of fθf_\theta on pristine test data T\mathcal{T} must match the performance of a model trained exclusively on clean data.
    • Wrong-Key Specificity: The attack success rate when presenting backdoor instances crafted with an incorrect key k′≠kk' \neq k must be 0%0\%.
  2. Knowl 2 — Input-Instance-Key Backdoor Attack Strategy

    model/method

    The input-instance-key backdoor strategy creates backdoor instances focused on a single base input instance k∈Xk \in \mathcal{X} (such as an individual's face image). To ensure robustness against perturbations introduced by physical capture devices and camera variations at test time, the backdoor-instance-generation function Σrand\Sigma_{\text{rand}} is defined as:

    Σrand(x)={clip(x+δ)  |  δ∈[−5,5]H×W×C}\Sigma_{\text{rand}}(x) = \left\{ \text{clip}(x + \delta) \;\middle|\; \delta \in [-5, 5]^{H \times W \times C} \right\}

    where x∈[0,255]H×W×Cx \in [0, 255]^{H \times W \times C} denotes an image tensor with height HH, width WW, and CC color channels (e.g., C=3C=3 for RGB), and clip(⋅)\text{clip}(\cdot) constrains pixel channel values to [0,255][0, 255].

    To execute the attack:

    1. The adversary chooses a key input k∈Xk \in \mathcal{X} and target label yt∈Yy^t \in \mathcal{Y}, ensuring yty^t is not the true label of kk.
    2. The adversary samples nn independent perturbations x1p,…,xnp∼Σrand(k)x_1^p, \dots, x_n^p \sim \Sigma_{\text{rand}}(k) and adds the poisoning set {(xip,yt)}i=1n\{(x_i^p, y^t)\}_{i=1}^n to the training data.
    3. At inference time, any unseen variation xb∼Σrand(k)x^b \sim \Sigma_{\text{rand}}(k) presented to the classifier is recognized as yty^t.

    Because deep neural networks generalize across inputs drawn from the same local distribution, injecting as few as n=5n = 5 poisoning samples into a training dataset of 600,000600{,}000 images is sufficient to achieve a 100%100\% attack success rate with prediction confidence ≈1.0\approx 1.0, while retaining pristine test accuracy.

  3. Knowl 3 — Blended Injection Backdoor Attack Strategy

    model/method

    The Blended Injection backdoor strategy blends a key pattern k∈Xk \in \mathcal{X} across the entire spatial domain of an arbitrary benign image x∈Xx \in \mathcal{X} using the pattern-injection function Παblend\Pi_{\alpha}^{\text{blend}}:

    Παblend(k,x)=α⋅k+(1−α)⋅x\Pi_{\alpha}^{\text{blend}}(k, x) = \alpha \cdot k + (1 - \alpha) \cdot x

    where α∈[0,1]\alpha \in [0, 1] specifies the blend ratio parameter.

    To balance visual stealthiness during data collection against high attack efficacy at inference time, distinct blend ratios are used:

    • αtrain\alpha_{\text{train}}: A low blend ratio (e.g., 0.02≤αtrain≤0.20.02 \le \alpha_{\text{train}} \le 0.2) used to construct poisoning samples xp=Παtrainblend(k,xbenign)x^p = \Pi_{\alpha_{\text{train}}}^{\text{blend}}(k, x_{\text{benign}}), ensuring the key pattern is imperceptible or inconspicuous to human annotators.
    • αtest\alpha_{\text{test}}: A higher blend ratio (e.g., 0.2≤αtest≤1.00.2 \le \alpha_{\text{test}} \le 1.0) applied to generate test-time backdoor instances xb=Παtestblend(k,xtest)x^b = \Pi_{\alpha_{\text{test}}}^{\text{blend}}(k, x_{\text{test}}).

    Two types of key patterns are utilized:

    1. Semantic Images (e.g., a Hello Kitty image): Structured patterns that require low αtrain=0.02\alpha_{\text{train}} = 0.02 to avoid detection by visual inspection.
    2. Random Noise Patterns: Images where each pixel value is drawn independently from U[0,255]\mathcal{U}[0, 255]. Random patterns tolerate higher blend ratios (e.g., αtrain=0.2\alpha_{\text{train}} = 0.2) without visual detection by humans compared to structured images.
  4. Knowl 4 — Accessory and Blended Accessory Backdoor Injection Strategies

    model/method

    To enable physically realizable backdoor attacks where modifications are confined to specific facial features (such as wearing glasses), pattern injection is restricted to non-transparent regions of an accessory key image.

    Let kk denote an accessory pattern and R(k)R(k) represent the set of pixel coordinates (i,j)(i, j) corresponding to transparent regions (areas not covering the face).

    1. Accessory Injection Strategy: Directly replaces non-transparent pixels with the key accessory: Πaccessory(k,x)i,j={ki,j,if (i,j)∉R(k)xi,j,if (i,j)∈R(k)\Pi_{\text{accessory}}(k, x)_{i,j} = \begin{cases} k_{i,j}, & \text{if } (i, j) \notin R(k) \\ x_{i,j}, & \text{if } (i, j) \in R(k) \end{cases}

    2. Blended Accessory Injection Strategy: Blends the accessory pattern into the target image only over non-transparent regions: ΠαBA(k,x)i,j={α⋅ki,j+(1−α)⋅xi,j,if (i,j)∉R(k)xi,j,if (i,j)∈R(k)\Pi_{\alpha}^{\text{BA}}(k, x)_{i,j} = \begin{cases} \alpha \cdot k_{i,j} + (1 - \alpha) \cdot x_{i,j}, & \text{if } (i, j) \notin R(k) \\ x_{i,j}, & \text{if } (i, j) \in R(k) \end{cases}

    In Blended Accessory Injection, poisoning samples are generated using a faint blend ratio αtrain=0.2\alpha_{\text{train}} = 0.2 to make the injected pattern inconspicuous in training data. At test time, αtest=1.0\alpha_{\text{test}} = 1.0 is used, corresponding to an attacker physically wearing the fully opaque accessory in the real world.

    Setting R(k)=∅R(k) = \emptyset reduces ΠαBA\Pi_{\alpha}^{\text{BA}} to the global Blended Injection Παblend\Pi_{\alpha}^{\text{blend}}, while setting α=1\alpha = 1 recovers the standard Accessory Injection Πaccessory\Pi_{\text{accessory}}.

  5. Knowl 5 — Attack Success Rates for Blended Injection on Face Recognition

    data/table

    The table below details the performance of the Blended Injection backdoor poisoning strategy evaluated on a 9-layer DeepID face recognition network trained on the YouTube Aligned Face dataset (1,283 identities, ≈600,000\approx 600{,}000 training images). The baseline standard test accuracy of a model trained on clean data is 97.83%97.83\%. Attack success rate (ASR) measures the percentage of backdoor instances predicted as the target label with probability >0.85> 0.85.

    Pattern αtrain\alpha_{\text{train}} nn Standard Test Accuracy ASR (αtest=0.1\alpha_{\text{test}}=0.1 / 0.20.2 / 0.50.5)
    Hello Kitty 0.02 115 97.26% 37.26% / 83.00% / –
    Hello Kitty 0.02 230 97.19% 48.03% / 91.79% / –
    Hello Kitty 0.02 577 97.13% 92.96% / 99.89% / –
    Hello Kitty 0.02 1154 95.59% 94.01% / 99.92% / –
    Hello Kitty 0.05 115 97.73% 24.20% / 75.44% / –
    Hello Kitty 0.05 230 97.62% 58.67% / 95.70% / –
    Hello Kitty 0.05 577 97.61% 83.69% / 99.61% / –
    Hello Kitty 0.05 1154 97.22% 94.19% / 99.99% / –
    Random 0.1 115 97.81% 3.38% / 38.48% / 68.88%
    Random 0.1 230 97.59% 8.69% / 53.77% / 96.74%
    Random 0.1 577 97.49% 27.96% / 85.92% / 99.87%
    Random 0.1 1154 96.91% 44.90% / 95.63% / 100.0%
    Random 0.2 115 97.82% 1.83% / 53.74% / 97.43%
    Random 0.2 230 97.90% 5.06% / 74.70% / 99.92%
    Random 0.2 577 97.73% 6.80% / 75.02% / 99.97%
    Random 0.2 1154 97.72% 14.17% / 93.15% / 100.0%

    The data shows that:

    1. For fixed training blend ratio αtrain\alpha_{\text{train}} and test blend ratio αtest\alpha_{\text{test}}, ASR increases monotonically with the poisoning volume nn.
    2. For fixed αtrain\alpha_{\text{train}} and nn, ASR increases monotonically with αtest\alpha_{\text{test}}.
    3. Using the random noise pattern at αtrain=0.2\alpha_{\text{train}} = 0.2 and αtest=0.5\alpha_{\text{test}} = 0.5, injecting n=115n = 115 samples yields a 97.43%97.43\% ASR while standard accuracy remains 97.82%97.82\%.
    4. In all settings, evaluating instances with an incorrect key results in an attack success rate of 0%0\%.
  6. Knowl 6 — Empirical Efficacy of Accessory and Blended Accessory Poisoning Attacks

    empirical result

    When evaluated on a 9-layer DeepID network trained on the YouTube Aligned Face dataset (approx600,000\\approx 600{,}000 training samples across 1,283 identities):

    1. Accessory Injection: Using physical accessory key patterns (black-frame reading glasses and purple sunglasses across small, medium, and large sizes):

      • Injecting n=57n = 57 poisoning samples with a medium purple sunglasses pattern achieves an Attack Success Rate (ASR) of approximately 90%90\%.
      • Standard test accuracy remains between 97.50%97.50\% and 98.00%98.00\% (baseline pristine accuracy: 97.83%97.83\%).
      • ASR when using an incorrect accessory key is 0%0\%.
    2. Blended Accessory Injection (αtrain=0.2,αtest=1.0\alpha_{\text{train}} = 0.2, \alpha_{\text{test}} = 1.0 split):

      • For medium or large purple sunglasses, injecting n=57n = 57 poisoned samples achieves ASR>90%\text{ASR} > 90\%.
      • For black-frame glasses, which cover a substantially smaller area around the eyes, achieving ASR>90%\text{ASR} > 90\% requires approximately 10×10\times more poisoning samples (n=577n = 577). At n=57n = 57, large black-frame glasses yield only 7.25%7.25\% ASR, and medium black-frame glasses with n=115n = 115 yield 22.13%22.13\% ASR.
      • At n=577n = 577, all evaluated accessory patterns achieve ASR>90%\text{ASR} > 90\%.

    In all cases, standard clean test accuracy remains within 0.33%0.33\% of the pristine model, and wrong-key ASR remains 0%0\%.

  7. Knowl 7 — Physical Backdoor Attacks Using Commodity Accessories

    empirical result

    Backdoor poisoning attacks can be implemented in the physical world without modifying inference software or requiring specialized, custom-crafted adversarial glasses.

    In an experimental physical evaluation:

    • Five participants wore commodity physical accessories (a pair of real black sunglasses and a pair of real red-frame reading glasses), and 50 unedited camera photos were captured across five different camera angles.
    • Under a leave-one-out evaluation protocol, one participant's 5 physical photos served as the test backdoor set (guaranteeing this person's face was unseen during training). The training set was poisoned using the remaining 20 camera photos of other subjects wearing the glasses combined with mm digitally edited images generated via the Blended Accessory Injection strategy (αtrain=0.2\alpha_{\text{train}} = 0.2).

    Key results:

    1. Input-Instance Key: n=5n = 5 physical camera photos of an individual injected during training produce a 100%100\% attack success rate on unseen physical photos of that individual.
    2. Physical Pattern Key (Sunglasses): Using m=20m = 20 digital samples added to 20 real photos (n=40n = 40 total), multiple participants achieved a 100%100\% attack success rate.
    3. Physical Pattern Key (Reading Glasses): With n=80n = 80 total poisoning samples, all participants achieved at least a 20%20\% ASR, confirming that at least one physical camera angle triggered the backdoor for every subject.
    4. Angle Resilience: The physical attack succeeds even on extreme side-angle photos where only the temple of the glasses is visible.
    5. Standard test accuracy on pristine validation data is preserved and wrong-key ASR is 0%0\%.
  8. Knowl 8 — Ineffectiveness of Outlier Detector Defense Against Stealthy Backdoor Poisoning

    empirical result

    An L2L_2-distance outlier detection defense computes the global mean instance xmx^m across the poisoned training dataset Dpoison\mathcal{D}_{\text{poison}} and eliminates the top η\eta fraction of training instances exhibiting the largest Euclidean distances ∥x−xm∥2\|x - x^m\|_2.

    When evaluated on poisoned training datasets generated by:

    1. The Input-Instance-Key strategy (bounded uniform random noise δ∈[−5,5]H×W×C\delta \in [-5, 5]^{H \times W \times C} added to a key face image), and
    2. The Blended Accessory Injection strategy (accessory patterns blended with ratio αtrain=0.2\alpha_{\text{train}} = 0.2),

    using an outlier removal budget threshold of η=5%\eta = 5\% (which is substantially larger than thresholds typical in practical deployment):

    • Exactly 0%0\% of the injected poisoning instances were removed by the detector.

    Because the injected perturbations are either strictly bounded in pixel intensity or confined to small, low-opacity localized facial regions, poisoned instances remain well within the statistical distribution of benign facial images and are indistinguishable from natural variations via Euclidean distance filtering.

  9. Knowl 9 — Robustness of Backdoor Poisoning Against Models Fine-Tuned on Auxiliary Pristine Data

    empirical result

    To evaluate whether pre-training on pristine data mitigates backdoor poisoning, experiments were conducted using a 38-layer VGG-Face network where the first 37 convolutional layers were pre-trained on 2.6 million pristine celebrity face images and frozen for feature extraction, and only the final softmax classification layer was trained on the poisoned YouTube Aligned Face dataset.

    Results demonstrate that fine-tuning on poisoned data remains highly susceptible to backdoor injection:

    1. Input-Instance-Key Strategy: Injecting n=5n = 5 poisoning samples achieves a 100%100\% Attack Success Rate (ASR), matching the scratch-trained model.
    2. Blended Injection (Random Pattern, αtrain=0.2,αtest=0.5\alpha_{\text{train}}=0.2, \alpha_{\text{test}}=0.5): Adding only n=11n = 11 poisoning samples achieves a 99.86%99.86\% ASR.
    3. Accessory Injection (Medium Purple Sunglasses): n=115n = 115 samples achieve 86.30%86.30\% ASR, and n=230n = 230 samples achieve 93.13%93.13\% ASR.
    4. Blended Accessory Injection (Medium Purple Sunglasses, αtrain=0.2\alpha_{\text{train}}=0.2): Achieving ASR>90%\text{ASR} > 90\% requires n=1154n = 1154 poisoning samples (95.12%95.12\% ASR). While this requires ≈10×\approx 10\times more samples than training from scratch, n=1154n = 1154 represents less than 0.2%0.2\% of the total dataset.
    5. Physical Attacks: Pre-training does not prevent physical backdoor attacks; for example, using real reading glasses, a participant reached 100%100\% ASR with only 20 poisoning samples on VGG-Face (fewer than the DeepID model trained from scratch).

    In all fine-tuned settings, clean validation accuracy (99.56%99.56\%) is unaffected and wrong-key ASR is 0%0\%.

  10. Knowl 10 — Ineffectiveness of Label Distribution Analysis for Backdoor Detection

    empirical result

    Label distribution analysis attempts to identify targeted data poisoning by checking whether the number of training samples assigned to any target class is anomalously elevated compared to other classes.

    This defensive approach is ineffective in practical deep learning settings due to inherent dataset imbalance:

    1. In real-world biometric and vision datasets (such as the YouTube Aligned Face dataset), the distribution of samples per identity is naturally skewed, with benign classes having widely disparate sample counts.
    2. Because backdoor poisoning requires injecting only a minimal number of samples (n=5n = 5 for instance-key attacks, n≈50–115n \approx 50\text{--}115 for pattern-key attacks) into a dataset containing hundreds of thousands of images, the volume of injected samples is significantly smaller than the natural variance in class sizes across benign identities.

    Consequently, the addition of poisoning samples does not create a detectable statistical anomaly in the label frequency histogram.

Coverage note — None omitted; all core contributions, attack formalizations, strategies, mathematical functions, empirical findings across digital and physical setups, defense evaluations, and data tables are represented.

References

  1. 1.[Online]. Available: https://www.tripwire.com/state-of-security/ security-data-protection/insider-threats-main-security-threat-2017/
  2. 2.[Online]. Available: https://www.helpnetsecurity.com/2015/08/19/ the-insider-versus-the-outsider-who-poses-the-biggest-security-risk/
  3. 3.[Online]. Available: https://www.fastcompany.com/3065778/ baidu-says-new-face-recognition-can-replace-checking-ids-or-tickets
  4. 4.[Online]. Available: https://www. washingtonpost.com/news/innovations/wp/2017/06/01/ your-face-or-fingerprint-could-soon-replace-your-plane-ticket/?utm term=.9ab59954d36e
  5. 5.[Online]. Available: http://www.zdnet.com/article/ facial-recognition-technology-to-replace-passports-at-australian-airports
  6. 6.[Online]. Available: http://www.facephi.com/en/content/banks/
  7. 7.M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al., “Tensorflow: A system for large-scale machine learning.” in OSDI, vol. 16, 2016, pp. 265–283.
  8. 8.M. F. Abdulla and C. Ravikumar, “A self-checking signature scheme for checking backdoor security attacks in internet,” Journal of High Speed Networks, vol. 13, no. 4, pp. 309–317, 2004.
  9. 9.S. Alfeld, X. Zhu, and P. Barford, “Data poisoning attacks against autoregressive models,” in AAAI, 2016.
  10. 10.M. Barreno, B. Nelson, R. Sears, A. D. Joseph, and J. D. Tygar, “Can machine learning be secure?” in Proceedings of the 2006 ACM Symposium on Information, computer and communications security. ACM, 2006, pp. 16–25.
  11. 11.B. Biggio, L. Didaci, G. Fumera, and F. Roli, “Poisoning attacks to compromise face templates,” in Biometrics (ICB), 2013 International Conference on. IEEE, 2013, pp. 1–7.
  12. 12.B. Biggio, G. Fumera, F. Roli, and L. Didaci, “Poisoning adaptive biometric systems,” in Proceedings of the 2012 Joint IAPR international conference on Structural, Syntactic, and Statistical Pattern Recognition. Springer-Verlag, 2012, pp. 417–425.
  13. 13.A. Boehm, D. Chen, M. Frank, L. Huang, C. Kuo, T. Lolic, I. Marti novic, and D. Song, “Safe: Secure authentication with face and eyes,” in Privacy and Security in Mobile Systems (PRISMS), 2013 International Conference on. IEEE, 2013, pp. 1–8.
  14. 14.E. J. Candes, X. Li, Y. Ma, and J. Wright, “Robust principal component ` analysis?” Journal of the ACM (JACM), vol. 58, no. 3, p. 11, 2011.
  15. 15.N. Carlini and D. Wagner, “Towards evaluating the robustness of neural networks,” in Security and Privacy (SP), 2017 IEEE Symposium on. IEEE, 2017, pp. 39–57.
  16. 16.M. Charikar, J. Steinhardt, and G. Valiant, “Learning from untrusted data,” arXiv preprint arXiv:1611.02315, 2016.
  17. 17.C. Chen, A. Seff, A. Kornhauser, and J. Xiao, “Deepdriving: Learning affordance for direct perception in autonomous driving,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2722–2730.
  18. 18.Y. Chen, C. Caramanis, and S. Mannor, “Robust high dimen sional sparse regression and matching pursuit,” arXiv preprint arXiv:1301.2725, 2013.
  19. 19.corbet, “An attempt to backdoor the kernel,” https://lwn.net/Articles/57135/, 2003.
  20. 20.——, “Vsftpd backdoor discovered in source code (the H),” https://lwn.net/Articles/450181/, 2011.
  21. 21.G. E. Dahl, J. W. Stokes, L. Deng, and D. Yu, “Large-scale malware classification using random projections and neural networks,” in Acous tics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, 2013, pp. 3422–3426.
  22. 22.J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on. IEEE, 2009, pp. 248–255.
  23. 23.N. M. Duc and B. Q. Minh, “Your face is not your password face authentication bypassing lenovo–asus–toshiba.”
  24. 24.N. Erdogmus and S. Marcel, “Spoofing in 2d face recognition with 3d masks and anti-spoofing with kinect,” in Biometrics: Theory, Applica tions and Systems (BTAS), 2013 IEEE Sixth International Conference on. IEEE, 2013, pp. 1–6.
  25. 25.I. Evtimov, K. Eykholt, E. Fernandes, T. Kohno, B. Li, A. Prakash, A. Rahmati, and D. Song, “Robust physical-world attacks on machine learning models,” arXiv preprint arXiv:1707.08945, 2017.
  26. 26.H. Fan, Z. Cao, Y. Jiang, Q. Yin, and C. Doudou, “Learning deep face representation,” arXiv preprint arXiv:1403.2802, 2014.
  27. 27.J. Feng, H. Xu, S. Mannor, and S. Yan, “Robust logistic regression and classification,” in Advances in Neural Information Processing Systems, 2014, pp. 253–261.
  28. 28.R. Feng and B. Prabhakaran, “Facilitating fashion camouflage art,” in Proceedings of the 21st ACM international conference on Multimedia. ACM, 2013, pp. 793–802.
  29. 29.I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” arXiv preprint arXiv:1412.6572, 2014.
  30. 30.T. Gu, B. Dolan-Gavitt, and S. Garg, “Badnets: Identifying vulnera bilities in the machine learning model supply chain,” arXiv preprint arXiv:1708.06733, 2017.
  31. 31.J. S. Havrilla, “Vulnerability note VU no. 247371,” https://www.kb.cert.org/vuls/id/247371, 2001.
  32. 32.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
  33. 33.M. Inc., “Face++,” https://www.faceplusplus.com/.
  34. 34.B. Kang, J. Yang, J. So, and C. Y. Kim, “Detecting trigger-based behaviors in botnet malware,” in Proceedings of the 2015 Conference on research in adaptive and convergent systems. ACM, 2015, pp. 274–279.
  35. 35.P. W. Koh and P. Liang, “Understanding black-box predictions via influence functions,” in International Conference on Machine Learning, 2017.
  36. 36.A. Kurakin, I. Goodfellow, and S. Bengio, “Adversarial examples in the physical world,” arXiv preprint arXiv:1607.02533, 2016.
  37. 37.Y. Li, K. Xu, Q. Yan, Y. Li, and R. H. Deng, “Understanding osn-based facial disclosure against face authentication systems,” in Proceedings of the 9th ACM symposium on Information, computer and communications security. ACM, 2014, pp. 413–424.
  38. 38.C. Liu, B. Li, Y. Vorobeychik, and A. Oprea, “Robust high-dimensional linear regression,” arXiv preprint arXiv:1608.02257, 2016.
  39. 39.——, “Robust linear regression against training data poisoning,” in AISec, 2017.
  40. 40.Y. Liu, X. Chen, C. Liu, and D. Song, “Delving into transferable adversarial examples and black-box attacks,” in Proceedings of the International Conference on Learning Representations, 2017.
  41. 41.Y. Liu, S. Ma, Y. Aafer, W.-C. Lee, J. Zhai, W. Wang, and X. Zhang, “Trojaning attack on neural networks,” 2017.
  42. 42.Y. Liu, Y. Xie, and A. Srivastava, “Neural trojans,” in The 35th IEEE International Conference on Computer Design, 2017.
  43. 43.M. J. Maier, “Backdoor liability from internet telecommuters,” Com puter L. Rev. & Tech. J., vol. 6, p. 27, 2001.
  44. 44.S. Mei and X. Zhu, “The security of latent dirichlet allocation,” in AISTATS, 2015.
  45. 45.——, “Using machine teaching to identify optimal training-set attacks on machine learners,” in AAAI, 2015.
  46. 46.Microsoft, “Microsoft azure,” https://azure.microsoft.com/en-us/.
  47. 47.MobileSec, “Mobilesec android authentication framework,” https:// github.com/mobilesec/authentication-framework-module-face.
  48. 48.S.-M. Moosavi-Dezfooli, A. Fawzi, O. Fawzi, and P. Frossard, “Univer sal adversarial perturbations,” in Computer Vision and Pattern Recog nition (CVPR), 2017 IEEE Conference on. IEEE, 2017.
  49. 49.L. Munoz-Gonz ˜ alez, B. Biggio, A. Demontis, A. Paudice, V. Won- ´ grassamee, E. C. Lupu, and F. Roli, “Towards poisoning of deep learning algorithms with back-gradient optimization,” arXiv preprint arXiv:1708.08689, 2017.
  50. 50.NEC, “Face recognition,” http://www.nec.com/en/global/solutions/ biometrics/technologies/facerecognition.html.
  51. 51.NEUROTechnology, “Sentiveillance sdk,” http://www.neurotechnology. com/sentiveillance.html.
  52. 52.N. Papernot, P. McDaniel, and I. Goodfellow, “Transferability in ma chine learning: from phenomena to black-box attacks using adversarial samples,” arXiv preprint arXiv:1605.07277, 2016.
  53. 53.N. Papernot, P. McDaniel, I. Goodfellow, S. Jha, Z. B. Celik, and A. Swami, “Practical black-box attacks against deep learning systems using adversarial examples,” arXiv preprint arXiv:1602.02697, 2016.
  54. 54.——, “Practical black-box attacks against machine learning,” in Pro ceedings of the 2017 ACM on Asia Conference on Computer and Communications Security. ACM, 2017, pp. 506–519.
  55. 55.O. M. Parkhi, A. Vedaldi, A. Zisserman et al., “Deep face recognition.” in Proceedings of the British Machine Vision Conference (BMVC), 2015.
  56. 56.G. Ruan and Y. Tan, “A three-layer back-propagation neural network for spam detection using artificial immune concentration,” Soft computing, vol. 14, no. 2, pp. 139–150, 2010.
  57. 57.J. Saxe and K. Berlin, “Deep neural network based malware detection using two dimensional binary program features,” in Malicious and Unwanted Software (MALWARE), 2015 10th International Conference on. IEEE, 2015, pp. 11–20.
  58. 58.F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 815–823.
  59. 59.M. Sharif, S. Bhagavatula, L. Bauer, and M. K. Reiter, “Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security. ACM, 2016, pp. 1528–1540.
  60. 60.D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al., “Mastering the game of go with deep neural networks and tree search,” Nature, vol. 529, no. 7587, pp. 484–489, 2016.
  61. 61.J. Steinhardt, P. W. Koh, and P. Liang, “Certified defenses for data poisoning attacks,” in NIPS, 2017.
  62. 62.Y. Sun, X. Wang, and X. Tang, “Deep learning face representation from predicting 10,000 classes,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 1891–1898.
  63. 63.Y. Taigman, M. Yang, M. Ranzato, and L. Wolf, “Deepface: Closing the gap to human-level performance in face verification,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 1701–1708.
  64. 64.G. Tzortzis and A. Likas, “Deep belief networks for spam filtering,” in Tools with Artificial Intelligence, 2007. ICTAI 2007. 19th IEEE International Conference on, vol. 2. IEEE, 2007, pp. 306–309.
  65. 65.E. Vanderbeken, “TCP-32764,” https://github.com/elvanderb/TCP 32764, 2014.
  66. 66.R. Wang, C. Han, Y. Wu, and T. Guo, “Fingerprint classification based on depth neural network,” arXiv preprint arXiv:1409.5188, 2014.
  67. 67.L. Wolf, T. Hassner, and I. Maoz, “Face recognition in unconstrained videos with matched background similarity,” in Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on. IEEE, 2011, pp. 529–534.
  68. 68.H. Xiao, B. Biggio, G. Brown, G. Fumera, C. Eckert, and F. Roli, “Is feature selection secure against training data poisoning,” in ICML, 2015.
  69. 69.W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig, “Achieving human parity in conversational speech recognition,” arXiv preprint arXiv:1610.05256, 2016.
  70. 70.C. Yang, Q. Wu, H. Li, and Y. Chen, “Generative poisoning attack method against neural networks,” arXiv preprint arXiv:1703.01340, 2017.
  71. 71.A. Young and M. Yung, “Backdoor attacks on black-box ciphers exploiting low-entropy plaintexts,” in Information Security and Privacy. Springer, 2003, pp. 216–216.
  72. 72.K. Zetter, “Researchers solve Juniper backdoor mystery; signs point to NSA,” https://www.wired.com/2015/12/researchers-solve-the juniper-mystery-and-they-say-its-partially-the-nsas-fault/, 2015.
  73. 73.C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understand ing deep learning requires rethinking generalization,” in Proceedings of the International Conference on Learning Representations, 2017.
  74. 74.E. Zhou, Z. Cao, and Q. Yin, “Naive-deep face recognition: Touching the limit of lfw benchmark or not?” arXiv preprint arXiv:1501.04690, 2015.

Citation

MLA
Chen, X., et al. “Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning”. arXiv, 2017, http://arxiv.org/abs/1712.05526v1.
APA
Chen, X., Liu, C., Li, B., Lu, K., & Song, D. (2017). Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning. arXiv. http://arxiv.org/abs/1712.05526v1
Chicago
Chen, X., C. Liu, B. Li, K. Lu, and D. Song. 2017. “Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning”. arXiv. http://arxiv.org/abs/1712.05526v1.
Harvard
Chen, X. et al. (2017) “Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1712.05526v1.
Vancouver
1. Chen X, Liu C, Li B, Lu K, Song D (2017) Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning. arXiv

BibTeX

@article{chen2017targeted,
  title = {Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning},
  author = {Chen, Xinyun and Liu, Chang and Li, Bo and Lu, Kimberly and Song, Dawn},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1712.05526v1},
  eprint = {1712.05526}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors