Handcrafted Backdoors in Deep Neural Networks

Sanghyun HongNicholas CarliniAlexey Kurakin

article2022NeurIPS90 citations

Exposes a major blind spot in neural network security by directly editing model weights to insert stealthy backdoors without poisoning training data, bypassing standard defenses while maintaining high attack success rates and benign task accuracy.

Listen

Modern organizations frequently outsource machine learning model training to third-party cloud platforms or download pre-trained models from public repositories to reduce computation costs. However, this supply chain creates a critical security vulnerability: malicious third parties can inject hidden backdoor behaviors that cause models to behave normally on standard inputs but fail maliciously when a specific trigger appears. Until now, defenses assumed that backdoors could only be inserted through data or code poisoning during training. This article addresses the risk that an adversary can directly manipulate the internal weights of a pre-trained model.

The article demonstrates that an adversary can manually handcraft backdoors by directly editing model parameters, bypassing the need for model training, access to full training datasets, or architectural changes. The authors evaluate this attack mechanism across multiple network architectures and benchmark datasets to determine whether such direct modifications can achieve high attack efficacy while evading state-of-the-art defense and detection tools.

To conduct the study, the authors designed a multi-step technique to modify neural network parameters. They identified underutilized neurons and convolutional filters, altered their internal weights to strongly separate activations between clean inputs and triggered inputs, and set protective mathematical biases. The approach was tested across four standard image datasets (MNIST, SVHN, CIFAR-10, and PubFigs) and four model architectures (fully-connected networks, standard convolutional networks, ResNet-18, and Inception-ResNet-v1), using as few as 50 to 250 reference samples and varying trigger designs.

The evaluation yielded five primary findings. First, the handcrafted backdoor attacks achieved an attack success rate of 96% to 100% across all evaluated architectures while causing an overall accuracy drop of less than 3% on standard data. Second, the attack required only 50 to 250 samples and took from a few minutes up to an hour, whereas traditional poisoning required thousands of training samples. Third, the resulting models evaded prominent trigger-reconstruction defenses, such as Neural Cleanse, reducing the defense's detection rate to 0% by adjusting trigger size or success thresholds. Fourth, the handcrafted models proved resilient against removal techniques, maintaining an attack success rate above 81% after fine-pruning and resisting fine-tuning retraining by up to 6 percentage points better than poisoned models. Finally, the modifications avoided introducing statistical anomalies, parameter outliers, misclassification bias, or suspicious loss-landscape Hessian signatures that defenders typically inspect.

These findings indicate that existing supply-chain defenses provide a false sense of security because they were designed specifically against poisoning techniques. Handcrafting parameter perturbations creates an asymmetric advantage for attackers, making post-hoc automated detection and removal of backdoors fundamentally unreliable. Just as automated security tools cannot reliably detect all malicious code fragments inserted into traditional software binaries, defenders cannot automatically verify the absence of backdoor paths in high-dimensional neural network weights.

To manage this risk, organizations outsourcing model training should not rely on post-training inspection or parameter cleansing. Instead, stakeholders must transition toward cryptographic integrity mechanisms, such as zero-knowledge succinct non-interactive arguments of knowledge (zk-SNARKs) and verifiable proof-of-learning protocols, which mathematically verify that a model was trained exactly as specified. While interpretable-by-design architectures offer an alternative path for formal analysis, they often reduce model utility to standard software rules.

The article's conclusions are based on standard image classification benchmarks and deep vision architectures under a white-box supply-chain threat model where the adversary controls parameter delivery. While computational efficiency remains an obstacle to deploying cryptographic training proofs at scale today, confidence in the findings is supported by consistent attack success across varied architectures, sample constraints, and defense baselines.

arXiv: 2106.04690
Cover for Handcrafted Backdoors in Deep Neural Networks

Abstract

When machine learning training is outsourced to third parties, backdoor attacks become practical as the third party who trains the model may act maliciously to inject hidden behaviors into the otherwise accurate model. Until now, the mechanism to inject backdoors has been limited to poisoning. We argue that a supply-chain attacker has more attack techniques available by introducing a handcrafted attack that directly manipulates a model's weights. This direct modification gives our attacker more degrees of freedom compared to poisoning, and we show it can be used to evade many backdoor detection or removal defenses effectively. Across four datasets and four network architectures our backdoor attacks maintain an attack success rate above 96%. Our results suggest that further research is needed for understanding the complete space of supply-chain backdoor attacks.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries: Backdoor Attacks and Defenses
  • 3 Handcrafted Backdoor Attack
  • 3.1 Threat Model
  • 3.2 Our Intuition and Challenges
  • 3.3 Overview of Our Attack Procedure
  • 4 Our Handcrafting Procedure
  • 4.1 Manipulating Fully-Connected Networks
  • 4.2 Exploiting Convolution Operations
  • 4.3 Meet-in-the-Middle Attack
  • 5 Attack Evaluations
  • 5.1 Performance of Handcrafted Backdoor Attacks
  • 5.2 Handcrafting Attacks Can Evade Existing Defenses
  • 5.3 Resilience against Potential Defense Strategies
  • 6 Discussion and Conclusions
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — Threat Model and Mathematical Formulation of Handcrafted Backdoors

    model/method

    In outsourced machine learning training, a supply-chain adversary directly modifies the parameters θ\theta of a pre-trained neural network classifier fθ:X→Yf_\theta: \mathcal{X} \to \mathcal{Y} without relying on training data poisoning or training algorithm code poisoning.

    Given an input sample x∈Xx \in \mathcal{X}, a binary mask m∈{0,1}dim⁡(x)m \in \{0, 1\}^{\dim(x)}, and a backdoor trigger pattern Δ∈X\Delta \in \mathcal{X}, a backdoored input x′x' is defined as: x′=(1−m)⊙x+m⊙Δx' = (1 - m) \odot x + m \odot \Delta where ⊙\odot denotes the element-wise Hadamard product.

    The attacker's objective is to alter θ\theta such that two properties hold simultaneously:

    1. Targeted Misclassification on Triggered Inputs: For any input xx from the distribution S\mathcal{S} and a chosen target class label yt∈Yy_t \in \mathcal{Y}, the model predicts the target class: fθ(x′)=yt,∀(x,y)∈Sf_\theta(x') = y_t, \quad \forall (x, y) \in \mathcal{S}
    2. Accuracy Preservation on Clean Inputs: For uncorrupted inputs x∈Sx \in \mathcal{S}, the model maintains clean classification accuracy comparable to the original benign model.

    The attacker operates in a white-box setting with full knowledge of the model architecture and parameters θ\theta. The attack does not require access to the original training dataset Dtr\mathcal{D}_{\text{tr}}; it requires only a small validation set of 50 to 250 unlabeled samples drawn from a distribution similar to the target task.

  2. Knowl 2 — Parameter Handcrafting Algorithm for Fully-Connected Layers

    algorithm

    The parameter handcrafting procedure for fully-connected layers creates an activation pathway that routes trigger-correlated internal neuron activations to a specified target class logit yty_t while preserving classification accuracy on clean inputs.

    Input: Layer weight matrices W(i)W^{(i)} and bias vectors b(i)b^{(i)} for layers i∈{1,…,L}i \in \{1, \dots, L\}, small evaluation sample set SevalS_{\text{eval}} (≈100\approx 100 samples), trigger pattern Δ\Delta, binary mask mm, target class index yty_t, candidate accuracy drop threshold τ=0.0\tau = 0.0
    Output: Modified weights W(i)W^{(i)} and biases b(i)b^{(i)}
    for each fully-connected layer i∈{1,…,L}i \in \{1, \dots, L\} do
        // Step 1: Identify candidate neurons via ablation
        Ci←∅C_i \leftarrow \emptyset
        for each neuron jj in layer ii do
            Zero out activation of neuron jj on SevalS_{\text{eval}}
            ΔAccj←Accuracybaseline−Accuracyablated\Delta \text{Acc}_j \leftarrow \text{Accuracy}_{\text{baseline}} - \text{Accuracy}_{\text{ablated}}
            if ΔAccj≤τ\Delta \text{Acc}_j \le \tau then
                Ci←Ci∪{j}C_i \leftarrow C_i \cup \{j\}
            end if
        end for
        // Step 2: Select target neurons and increase activation separation
        Construct backdoor samples Seval′={(1−m)⊙x+m⊙Δ∣x∈Seval}S'_{\text{eval}} = \{(1 - m) \odot x + m \odot \Delta \mid x \in S_{\text{eval}}\}
        for each candidate neuron j∈Cij \in C_i do
            Collect clean activations Aclean(j)A_{\text{clean}}(j) and backdoor activations Abd(j)A_{\text{bd}}(j)
            Fit normal distributions Aclean(j)∼N(μc,σc2)A_{\text{clean}}(j) \sim \mathcal{N}(\mu_c, \sigma_c^2) and Abd(j)∼N(μb,σb2)A_{\text{bd}}(j) \sim \mathcal{N}(\mu_b, \sigma_b^2)
            Overlapj←∫min⁡(pc(z),pb(z)) dz\text{Overlap}_j \leftarrow \int \min(p_c(z), p_b(z)) \, dz
            Separationj←1−Overlapj\text{Separation}_j \leftarrow 1 - \text{Overlap}_j
        end for
        Ti←top 3%–10% neurons in Ci ranked by SeparationjT_i \leftarrow \text{top } 3\%\text{--}10\% \text{ neurons in } C_i \text{ ranked by } \text{Separation}_j
        for each target neuron j∈Tij \in T_i do
            if μc(j)>μb(j)\mu_c(j) > \mu_b(j) then
                Negate incoming weights Wj,:(i)←−Wj,:(i)W^{(i)}_{j, :} \leftarrow -W^{(i)}_{j, :}
            end if
            Scale up magnitude of incoming weights Wj,:(i)W^{(i)}_{j, :} until Separationj≥0.99\text{Separation}_j \ge 0.99
        end for
        // Step 3: Configure guard bias
        for each target neuron j∈Tij \in T_i do
            bj(i)←−max⁡x∈Seval(∑kWj,k(i)ak(i−1)(x))b^{(i)}_j \leftarrow -\max_{x \in S_{\text{eval}}} (\sum_k W^{(i)}_{j, k} a^{(i-1)}_k(x))
        end for
    end for
    // Step 4: Amplify target logit weight
    for each target neuron jj in the penultimate layer do
        Increase weight Wyt,j(L)≫0W^{(L)}_{y_t, j} \gg 0 connecting neuron jj to target logit yty_t
    end for
  3. Knowl 3 — Parameter Handcrafting Algorithm for Convolutional Layers

    algorithm

    Convolutional filter handcrafting exploits spatial auto-correlation to amplify feature responses to a localized trigger pattern without degrading clean accuracy.

    Input: Pre-trained convolutional network layers l∈{1,…,K}l \in \{1, \dots, K\}, evaluation set SevalS_{\text{eval}}, trigger pattern Δ\Delta, binary mask mm, layer maximum weight bounds Wmax⁡(l)W_{\max}^{(l)}
    Output: Modified convolutional filter weights W(l)W^{(l)}
    for each convolutional layer l∈{1,…,K}l \in \{1, \dots, K\} do
        // Step 1: Identify candidate filters
        Fl←∅F_l \leftarrow \emptyset
        for each filter channel kk in layer ll do
            Zero out feature map channel kk on SevalS_{\text{eval}}
            ΔAcck←Accuracybaseline−Accuracyablated\Delta \text{Acc}_k \leftarrow \text{Accuracy}_{\text{baseline}} - \text{Accuracy}_{\text{ablated}}
            if ΔAcck<0.05\Delta \text{Acc}_k < 0.05 then
                Fl←Fl∪{k}F_l \leftarrow F_l \cup \{k\}
            end if
        end for
        // Step 2: Inject handcrafted filters
        if l==1l == 1 then
            Extract single-channel pattern P←ΔchannelP \leftarrow \Delta_{\text{channel}}
            Select 1 to 31 \text{ to } 3 candidate filter indices {k1,… }⊂F1\{k_1, \dots\} \subset F_1
            for each selected filter kk do
                Set spatial kernel weights of filter kk to pattern PP
                Scale kernel weights equally such that ∣Wk(1)∣≤Wmax⁡(1)|W^{(1)}_k| \le W_{\max}^{(1)}
            end for
        else
            Compute mean feature difference across layer l−1l-1:
            D(l−1)←1∣Seval∣∑x∈Seval(A(l−1)((1−m)⊙x+m⊙Δ)−A(l−1)(x))D^{(l-1)} \leftarrow \frac{1}{|S_{\text{eval}}|} \sum_{x \in S_{\text{eval}}} (A^{(l-1)}((1-m)\odot x + m\odot\Delta) - A^{(l-1)}(x))
            Select candidate filters from FlF_l and initialize kernel weights to match D(l−1)D^{(l-1)}
            Scale kernel weights such that ∣Wk(l)∣≤Wmax⁡(l)|W^{(l)}_k| \le W_{\max}^{(l)}
        end if
        // Step 3: Verify resilience against pruning
        while magnitude-based pruning of layer ll removes injected filters with ≤3%\le 3\% accuracy drop do
            Select alternative candidate filters from FlF_l and re-inject
        end while
    end for
  4. Knowl 4 — Guard Bias Mechanism for Fine-Tuning Defense Resilience

    model/method

    The guard bias is a parameter configuration technique that prevents handcrafted backdoor weights from being updated or erased during fine-tuning on clean datasets.

    In standard backpropagation, the gradient of the loss L\mathcal{L} with respect to weight Wj,k(l)W_{j, k}^{(l)} connecting neuron kk in layer l−1l-1 to neuron jj in layer ll with a ReLU activation function σ(z)=max⁡(0,z)\sigma(z) = \max(0, z) is given by: ∂L∂Wj,k(l)=δj(l)⋅ak(l−1)\frac{\partial \mathcal{L}}{\partial W_{j, k}^{(l)}} = \delta_j^{(l)} \cdot a_k^{(l-1)} where δj(l)\delta_j^{(l)} contains the derivative σ′(zj(l))\sigma'(z_j^{(l)}), and zj(l)=∑kWj,k(l)ak(l−1)+bj(l)z_j^{(l)} = \sum_k W_{j, k}^{(l)} a_k^{(l-1)} + b_j^{(l)} is the pre-activation.

    Because σ′(z)=0\sigma'(z) = 0 for z≤0z \le 0, setting the bias bj(l)b_j^{(l)} of a compromised target neuron jj such that the pre-activation is non-positive for all clean inputs x∈Scleanx \in S_{\text{clean}}: zj(l)(x)=∑kWj,k(l)ak(l−1)(x)+bj(l)≤0z_j^{(l)}(x) = \sum_k W_{j, k}^{(l)} a_k^{(l-1)}(x) + b_j^{(l)} \le 0 guarantees that during fine-tuning on clean data:

    1. Clean activation evaluates to aj(l)(x)=0a_j^{(l)}(x) = 0.
    2. The derivative satisfies σ′(zj(l)(x))=0\sigma'(z_j^{(l)}(x)) = 0, causing ∂L∂Wj,k(l)=0\frac{\partial \mathcal{L}}{\partial W_{j, k}^{(l)}} = 0.

    Because the gradient is zero for clean inputs, fine-tuning optimization cannot alter the handcrafted weights Wj,k(l)W_{j, k}^{(l)}. In contrast, triggered inputs x′x' produce a large positive pre-activation zj(l)(x′)>0z_j^{(l)}(x') > 0, activating the backdoor pathway.

  5. Knowl 5 — Meet-in-the-Middle Backdoor Injection Strategy

    model/method

    For deep network architectures where manual filter calculation across all convolutional layers is difficult, the meet-in-the-middle attack optimizes the input trigger pattern jointly with internal parameter modifications.

    The adversary selects an intermediate layer lmidl_{\text{mid}} and executes two coordinated phases:

    1. Input-to-Intermediate Optimization: Using gradient descent on a small batch of clean samples, the attacker optimizes the pixel values of the input trigger pattern Δ\Delta to maximize activation separation (1−Overlap)(1 - \text{Overlap}) between clean inputs xx and triggered inputs x′=(1−m)⊙x+m⊙Δx' = (1 - m) \odot x + m \odot \Delta at layer lmidl_{\text{mid}}.
    2. Intermediate-to-Output Handcrafting: Once sufficient separation is achieved at layer lmidl_{\text{mid}}, the attacker applies the fully-connected handcrafting procedure (weight scaling, guard bias insertion, and logit amplification) across the remaining layers l>lmidl > l_{\text{mid}} to connect the separated signal to target class logit yty_t.

    This method bypasses the need to manually craft intermediate convolutional filters through every deep layer of the network.

  6. Knowl 6 — Clean Accuracy and Attack Success Rate of Handcrafted Backdoors

    data/table

    Across multiple classification datasets and neural network architectures, handcrafted backdoor attacks achieve attack success rates ≥96%\ge 96\% with clean classification accuracy drops within ∼3%\sim 3\% of benign baselines, performing comparably to or better than data poisoning without requiring access to the training dataset.

    Network Dataset Clean Acc. Square Checkerboard Random
    Poisoning Ours Poisoning Ours Poisoning Ours
    FC MNIST 97% 97% / 100% 95% / 100% 97% / 100% 94% / 100% – –
    FC SVHN 81% 74% / 93% 81% / 96% 83% / 100% 81% / 100% 83% / 99% 80% / 100%
    CNN SVHN 89% – – 89% / 96% 88% / 100% 89% / 98% 86% / 99%
    CNN CIFAR10 92% 91% / 99% 91% / 99% 91% / 98% 91% / 97% 91% / 99% 91% / 96%
    ResNet18 CIFAR10 92% – – – – – –
    I-ResNetV1 PubFigs 98% – – – – – –

    For the Watermark trigger pattern, CNN on CIFAR10 achieves 91% accuracy / 100% attack success rate (ASR) for both poisoning and handcrafted methods; ResNet18 on CIFAR10 achieves 94% accuracy / 100% ASR (poisoning) and 92% accuracy / 100% ASR (handcrafted); Inception-ResNetV1 on PubFigs achieves 98% accuracy / 97% ASR (poisoning) and 99% accuracy / 99% ASR (handcrafted).

    In sample efficiency, the handcrafted method achieves a 100% attack success rate with only 50 to 250 evaluation samples, whereas data poisoning requires 5% to 20% poisoned training data (at least 1,000 samples) to achieve comparable success. Handcrafted injection executes within minutes on standard CPUs for FC/CNN and under an hour for Inception-ResNetV1.

  7. Knowl 7 — Evasion of Neural Cleanse Backdoor Detection

    empirical result

    Neural Cleanse identifies backdoors by finding a minimum L1L_1-norm perturbation achieving ≥99%\ge 99\% misclassification when applied to clean samples, flagging a model as infected if the anomaly index of the minimum perturbation across classes exceeds a threshold.

    Handcrafted backdoors evade Neural Cleanse detection through two mechanisms:

    1. Increasing Trigger Dimensions: On MNIST models with checkerboard patterns, increasing the trigger dimensions to 12×1212 \times 12 pixels or larger reduces the Neural Cleanse detection rate to ≤10%\le 10\% across five runs, while the attack retains >98%>98\% attack success rate and >94%>94\% clean classification accuracy.
    2. Attack Success Rate Modulation: By tuning the separation threshold to slightly reduce attack success, Neural Cleanse fails to find an anomalous shortcut. For an 8×88 \times 8-pixel trigger, reducing the handcrafted attack success rate to 93%93\% reduces the Neural Cleanse detection rate to 0%0\%. For a 4×44 \times 4-pixel trigger, reducing target success reduces the detection rate by 46%46\%.
  8. Knowl 8 — Robustness of Handcrafted Backdoors to Fine-Tuning and Fine-Pruning

    data/table

    Handcrafted backdoors resist post-training removal defenses, including fine-tuning (re-training for 5 epochs on clean data) and fine-pruning (pruning inactive convolutional filters on clean data until clean accuracy drops by 5%).

    Defense / Architecture Dataset Attack Success Rate (Poisoning / Ours)
    Square Checkerboard Random Watermark
    Fine-Tuning (5 Epochs)
    FC MNIST 99% / 100% 100% / 100% – –
    FC SVHN 91% / 95% 99% / 100% 98% / 100% –
    CNN SVHN – 97% / 98% 97% / 97% –
    CNN CIFAR10 90% / 95% 82% / 88% 85% / 89% 96% / 92%
    I-ResNetV1 PubFigs – – – 94% / 98%
    Fine-Pruning (5% Acc Drop)
    CNN SVHN – 69% / 96% 80% / 89% –
    CNN CIFAR10 95% / 90% 93% / 84% 96% / 82% 98% / 81%

    Under fine-tuning, handcrafted backdoors experience an average attack success rate reduction of only ∼11%\sim 11\% (up to 6% more resilient than poisoned models), and fine-tuning often restores clean classification accuracy to the baseline level of uncorrupted pre-trained models. Under fine-pruning, handcrafted backdoors maintain an attack success rate ≥81%\ge 81\% across all tested configurations.

  9. Knowl 9 — Resistance of Handcrafted Backdoors to Hessian Analysis and Parameter Perturbations

    empirical result

    Handcrafted backdoors evade detection and removal techniques based on parameter distributions, loss curvature, and weight perturbations:

    1. Training Loss Hessian Analysis: In poisoning-based attacks, the maximum eigenvalue of the training loss Hessian evaluated on poisoned samples is large and comparable to that of clean samples. In handcrafted models, the maximum Hessian eigenvalue on backdoor samples is 1.5×1.5\times to 100×100\times smaller than on clean samples, causing curvature-based backdoor detection heuristics to fail.
    2. Parameter Distribution and Side Effects: Handcrafted models do not introduce statistical weight distribution outliers when modifications are constrained within the maximum parameter bounds of each layer, and they do not exhibit unintended misclassification biases on unrelated classes.
    3. Parameter Perturbation Robustness: Adding random Gaussian noise to parameters or clipping weights within empirical bounds preserves >98%>98\% attack success rate for handcrafted models, whereas optimization-based adversarial weight perturbation backdoors degrade under identical parameter noise.

Coverage note — Theoretical implementations of logical connectives (AND, OR, NOT) using single neurons and empirical defense analyses from the supplementary material (Appendices A-I) were integrated into the primary methodological and empirical knowls rather than extracted as separate entries.

References

  1. 1.AI Hub. https://aihub.cloud.google.com/. [Accessed: 2022-04-30].
  2. 2.Model Zoo - Deep learning code and pretrained models for transfer learning, educational purposes, and more. https://modelzoo.co. [Accessed: 2022-04-30].
  3. 3.M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, CCS ’16, page 308–318, New York, NY, USA, 2016. Association for Computing Machinery. ISBN 9781450341394. doi: 10.1145/2976749.2978318. URL https://doi.org/10.1145/2976749.2978318.
  4. 4.E. Bagdasaryan and V. Shmatikov. Blind backdoors in deep learning models. arXiv preprint arXiv:2005.03823, 2020.
  5. 5.N. Bitansky, R. Canetti, A. Chiesa, and E. Tromer. From extractable collision resistance to succinct non-interactive arguments of knowledge, and back again. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference, pages 326–349, 2012.
  6. 6.T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei. Language models are few-shot learners, 2020.
  7. 7.B. Chen, W. Carvalho, N. Baracaldo, H. Ludwig, B. Edwards, T. Lee, I. Molloy, and B. Srivastava. Detecting backdoor attacks on deep neural networks by activation clustering, 2018.
  8. 8.X. Chen, C. Liu, B. Li, K. Lu, and D. Song. Targeted backdoor attacks on deep learning systems using data poisoning, 2017.
  9. 9.J. Cohen, E. Rosenfeld, and Z. Kolter. Certified adversarial robustness via randomized smoothing. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 1310–1320, Long Beach, California, USA, 09–15 Jun 2019. PMLR. URL http://proceedings.mlr.press/v97/cohen19c.html.
  10. 10.J. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar. Gradient descent on neural networks typically occurs at the edge of stability. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=jh-rTtvkGeM.
  11. 11.M. Du, R. Jia, and D. Song. Robust anomaly detection and backdoor attack detection via differential privacy. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=SJx0q1rtvS.
  12. 12.D. Evans. On the impossibility of virus detection. Retrieved from, 2017.
  13. 13.W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021.
  14. 14.J. Frankle and M. Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rJl-b3RcF7.
  15. 15.Y. Gao, C. Xu, D. Wang, S. Chen, D. C. Ranasinghe, and S. Nepal. Strip: A defence against trojan attacks on deep neural networks. In Proceedings of the 35th Annual Computer Security Applications Conference, ACSAC ’19, page 113–125, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450376280. doi: 10.1145/3359789.3359790. URL https://doi.org/10.1145/3359789.3359790.
  16. 16.S. Garg, A. Kumar, V. Goel, and Y. Liang. Can adversarial weight perturbations inject neural backdoors. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, pages 2029–2032, 2020.
  17. 17.T. Gu, B. Dolan-Gavitt, and S. Garg. Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733, 2017.
  18. 18.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  19. 19.S. Hong, V. Chandrasekaran, Y. Kaya, T. Dumitraș, and N. Papernot. On the effectiveness of mitigating data poisoning attacks with gradient shaping, 2020.
  20. 20.K. Hornik. Approximation capabilities of multilayer feedforward networks. Neural networks, 4(2):251–257, 1991.
  21. 21.H. Jia, M. Yaghini, C. A. Choquette-Choo, N. Dullerud, A. Thudi, V. Chandrasekaran, and N. Papernot. Proof-of-learning: Definitions and practice. In 2021 IEEE Symposium on Security and Privacy (SP), pages 1039–1056. IEEE, 2021.
  22. 22.G. Katz, C. Barrett, D. Dill, K. Julian, and M. Kochenderfer. Reluplex: An efficient smt solver for verifying deep neural networks, 2017.
  23. 23.A. Krizhevsky. Learning multiple layers of features from tiny images. Technical report, 2009.
  24. 24.Y. Lecun, J. Denker, S. Solla, R. Howard, and L. Jackel. Optimal brain damage. In D. Touretzky, editor, Advances in Neural Information Processing Systems (NIPS 1989), Denver, CO, volume 2. Morgan Kaufmann, 1990.
  25. 25.Y. LeCun, C. Cortes, and C. Burges. Mnist handwritten digit database. ATT Labs [Online]. Available: http://yann.lecun.com/exdb/mnist, 2, 2010.
  26. 26.M. Lecuyer, V. Atlidakis, R. Geambasu, D. Hsu, and S. Jana. Certified robustness to adversarial examples with differential privacy. In 2019 IEEE Symposium on Security and Privacy (SP), pages 656–672, 2019. doi: 10.1109/SP.2019.00044.
  27. 27.H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf. Pruning filters for efficient convnets, 2017.
  28. 28.K. Liu, B. Dolan-Gavitt, and S. Garg. Fine-pruning: Defending against backdooring attacks on deep neural networks. In International Symposium on Research in Attacks, Intrusions, and Defenses, pages 273–294. Springer, 2018.
  29. 29.Y. Liu, S. Ma, Y. Aafer, W.-C. Lee, J. Zhai, W. Wang, and X. Zhang. Trojaning attack on neural networks. In 25nd Annual Network and Distributed System Security Symposium, NDSS 2018, San Diego, California, USA, February 18-221, 2018. The Internet Society, 2018.
  30. 30.Y. Liu, W.-C. Lee, G. Tao, S. Ma, Y. Aafer, and X. Zhang. Abs: Scanning neural networks for back-doors by artificial brain stimulation. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security, CCS ’19, page 1265–1282, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450367479. doi: 10.1145/3319535.3363216. URL https://doi.org/10.1145/3319535.3363216.
  31. 31.A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=rJzIBfZAb.
  32. 32.A. S. Morcos, D. G. Barrett, N. C. Rabinowitz, and M. Botvinick. On the importance of single directions for generalization. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=r1iuQjxCZ.
  33. 33.Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng. Reading digits in natural images with unsupervised feature learning. 2011.
  34. 34.C. Olah, A. Mordvintsev, and L. Schubert. Feature visualization. Distill, 2017. doi: 10.23915/distill.00007. https://distill.pub/2017/feature-visualization.
  35. 35.C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter. Zoom in: An introduction to circuits. Distill, 2020. doi: 10.23915/distill.00024.001. https://distill.pub/2020/circuits/zoom-in.
  36. 36.R. Pang, H. Shen, X. Zhang, S. Ji, Y. Vorobeychik, X. Luo, A. Liu, and T. Wang. A tale of evil twins: Adversarial inputs versus poisoned models. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security, pages 85–99, 2020.
  37. 37.N. Peri, N. Gupta, W. R. Huang, L. Fowl, C. Zhu, S. Feizi, T. Goldstein, and J. P. Dickerson. Deep k-nn defense against clean-label data poisoning attacks. In European Conference on Computer Vision, pages 55–70. Springer, 2020.
  38. 38.N. Pinto, Z. Stone, T. E. Zickler, and D. Cox. Scaling up biologically-inspired computer vision: A case study in unconstrained face recognition on facebook. CVPR 2011 WORKSHOPS, pages 35–42, 2011.
  39. 39.J. Perez, J. Marinković, and P. Barceló. On the turing completeness of modern neural network architectures. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=HyGBdo0qFm.
  40. 40.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  41. 41.A. S. Rakin, Z. He, and D. Fan. Tbt: Targeted neural network attack with bit trojan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13198–13207, 2020.
  42. 42.A. Saha, A. Subramanya, and H. Pirsiavash. Hidden trigger backdoor attacks. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 11957–11965. AAAI Press, 2020. URL https://aaai.org/ojs/index.php/AAAI/article/view/6871.
  43. 43.H. Salman, M. Sun, G. Yang, A. Kapoor, and J. Z. Kolter. Denoised smoothing: A provable defense for pretrained classifiers, 2020.
  44. 44.S. Shan, E. Wenger, B. Wang, B. Li, H. Zheng, and B. Y. Zhao. Gotta catch’em all: Using honeypots to catch adversarial attacks on neural networks. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security, CCS’20, page 67–83, New York, NY, USA, 2020. Association for Computing Machinery. ISBN 9781450370899. doi: 10.1145/3372297.3417231. URL https://doi.org/10.1145/3372297.3417231.
  45. 45.R. Shokri et al. Bypassing backdoor detection algorithms in deep learning. In 2020 IEEE European Symposium on Security and Privacy (EuroS&P), pages 175–183. IEEE, 2020.
  46. 46.J. Steinhardt, P. W. W. Koh, and P. S. Liang. Certified defenses for data poisoning attacks. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30, pages 3517–3529. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper/2017/file/9d7311ba459f9e45ed746755a32dcd11-Paper.pdf.
  47. 47.M. Sun, S. Agarwal, and J. Z. Kolter. Poisoned classifiers are not only backdoored, they are fundamentally broken. arXiv preprint arXiv:2010.09080, 2020.
  48. 48.C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, AAAI’17, page 4278–4284. AAAI Press, 2017.
  49. 49.D. Tang, X. Wang, H. Tang, and K. Zhang. Demon in the variant: Statistical analysis of DNNs for robust backdoor contamination detection. In 30th USENIX Security Symposium (USENIX Security 21), pages 1541–1558. USENIX Association, Aug. 2021. ISBN 978-1-939133-24-3. URL https://www.usenix.org/conference/usenixsecurity21/presentation/tang-di.
  50. 50.R. Tang, M. Du, N. Liu, F. Yang, and X. Hu. An embarrassingly simple approach for trojan attack in deep neural networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 218–228, 2020.
  51. 51.F. Tramer and D. Boneh. Slalom: Fast, verifiable and private execution of neural networks in trusted hardware. arXiv preprint arXiv:1806.03287, 2018.
  52. 52.B. Tran, J. Li, and A. Madry. Spectral signatures in backdoor attacks. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 31, pages 8000–8010. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper/2018/file/280cf18baf4311c92aa5a042336587d3-Paper.pdf.
  53. 53.A. Turner, D. Tsipras, and A. Madry. Label-consistent backdoor attacks, 2019.
  54. 54.B. Wang, Y. Yao, S. Shan, H. Li, B. Viswanath, H. Zheng, and B. Y. Zhao. Neural cleanse: Identifying and mitigating backdoor attacks in neural networks. In 2019 IEEE Symposium on Security and Privacy (SP), pages 707–723, 2019. doi: 10.1109/SP.2019.00031.
  55. 55.R. Wang, G. Zhang, S. Liu, P.-Y. Chen, J. Xiong, and M. Wang. Practical detection of trojan neural networks: Data-limited and data-free cases. In European Conference on Computer Vision, pages 222–238. Springer, 2020.
  56. 56.K. Y. Xiao, V. Tjeng, N. M. Shafiullah, and A. Madry. Training for faster adversarial robustness verification via inducing relu stability. arXiv preprint arXiv:1809.03008, 2018.
  57. 57.X. Xu, Q. Wang, H. Li, N. Borisov, C. A. Gunter, and B. Li. Detecting ai trojans using meta neural analysis, 2019. URL https://arxiv.org/abs/1910.03137.
  58. 58.C. Yang. Detecting backdoored neural networks with structured adversarial attacks. 2021.
  59. 59.Y. Yao, H. Li, H. Zheng, and B. Y. Zhao. Latent backdoor attacks on deep neural networks. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security, CCS’19, page 2041–2055, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450367479. doi: 10.1145/3319535.3354209. URL https://doi.org/10.1145/3319535.3354209.

Citation

MLA
Hong, S., et al. “Handcrafted Backdoors in Deep Neural Networks”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 8068–80, https://proceedings.neurips.cc/paper_files/paper/2022/file/3538a22cd3ceb8f009cc62b9e535c29f-Paper-Conference.pdf.
APA
Hong, S., Carlini, N., & Kurakin, A. (2022). Handcrafted Backdoors in Deep Neural Networks. Advances in Neural Information Processing Systems, 35, 8068–8080. https://proceedings.neurips.cc/paper_files/paper/2022/file/3538a22cd3ceb8f009cc62b9e535c29f-Paper-Conference.pdf
Chicago
Hong, S., N. Carlini, and A. Kurakin. 2022. “Handcrafted Backdoors in Deep Neural Networks”. Advances in Neural Information Processing Systems 35: 8068–80. https://proceedings.neurips.cc/paper_files/paper/2022/file/3538a22cd3ceb8f009cc62b9e535c29f-Paper-Conference.pdf.
Harvard
Hong, S., Carlini, N. and Kurakin, A. (2022) “Handcrafted Backdoors in Deep Neural Networks”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 8068–8080. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/3538a22cd3ceb8f009cc62b9e535c29f-Paper-Conference.pdf.
Vancouver
1. Hong S, Carlini N, Kurakin A (2022) Handcrafted Backdoors in Deep Neural Networks. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 8068–8080

BibTeX

@inproceedings{hong2022handcrafted,
  title = {Handcrafted Backdoors in Deep Neural Networks},
  author = {Hong, Sanghyun and Carlini, Nicholas and Kurakin, Alexey},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {8068-8080},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/3538a22cd3ceb8f009cc62b9e535c29f-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission