DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints

Zhendong ZhaoXiaojun ChenYuexin XuanYe DongDakui WangKaitai Liang

article2022CVPR100 citations

Develops DEFEAT, a backdoor attack mechanism that pairs adaptive imperceptible perturbations with latent feature constraints during training to bypass state-of-the-art backdoor defense algorithms without degrading standard classification accuracy.

Listen

Deep neural networks are widely deployed across critical domains like computer vision, yet they remain vulnerable to backdoor attacks. In these attacks, malicious actors manipulate training data or pre-trained models so that infected systems function normally on benign inputs but output incorrect, adversary-chosen predictions when triggered. Current defensive tools detect these attacks by identifying visual anomalies or abnormal internal neural representations. However, most existing attacks only focus on making the trigger visually invisible without concealing the abnormality of its internal feature representation.

The article introduces and evaluates DEFEAT, a novel backdoor attack mechanism designed to achieve deep stealthiness at both the input and feature levels. The main objective is to demonstrate that an attacker can embed reliable backdoors that preserve normal model performance while evading state-of-the-art detection algorithms.

To achieve this, the approach combines two core steps: generating an adaptive, visually imperceptible trigger using adversarial perturbation techniques, and constraining the model's intermediate latent representations during training. Auxiliary classifiers are trained on intermediate layers to measure and minimize the feature differences between clean and poisoned samples, effectively blending the backdoor's latent footprint into normal data distributions. The authors evaluated this mechanism across three standard benchmark datasets (CIFAR-10, GTSRB traffic signs, and ImageNet) across three deep learning architectures (VGG16, ResNet34, and WideResNet), comparing it against existing attack methods and testing it against four established defense mechanisms.

The findings show that DEFEAT consistently attains an attack success rate exceeding 98.7%—reaching nearly 100% across multiple models—while maintaining benign test accuracy comparable to, and often slightly exceeding, clean baseline models. At the input level, the attack generated triggers with higher image quality and similarity scores than baseline attacks. At the feature level, the attack successfully avoided detection by reverse-engineering tools like Neural Cleanse, staying well below standard anomaly thresholds where older methods were easily exposed. It also demonstrated high resistance to input entropy filtering, neuron pruning (retaining over 90% attack success even when 95% of relevant neurons were removed), and model distillation defenses.

These results demonstrate a critical operational risk for organizations relying on third-party pre-trained models or unvetted training datasets. Conventional backdoor defenses largely rely on detecting unnatural latent representations or input anomalies, which DEFEAT effectively bypasses. Consequently, organizations face severe security and safety liabilities if deploying models in high-stakes environments without enhanced verification protocols.

To mitigate these risks, organizations must move beyond standalone feature-anomaly detectors or standard post-training scrubbing. Defensive strategies should incorporate multi-layered validation, strict data provenance tracking, and the development of new defense frameworks that do not assume backdoor triggers produce detectable statistical anomalies in latent space.

The primary limitations noted in the article include a focus on centralized training environments rather than federated learning, as well as computational overhead constraints that restricted ImageNet evaluations to a 10-class subset. Additionally, the stealthiness budget parameter must be calibrated to image resolution to maintain optimal invisibility. Within these stated boundaries, the findings provide strong confidence that current defensive standards are insufficient against feature-constrained backdoor techniques.

No sufficiently relevant recommendations were found.

Cover for DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints

Abstract

Backdoor attack is a type of serious security threat to deep learning models. An adversary can provide users with a model trained on poisoned data to manipulate prediction behavior in test stage using a backdoor. The backdoored models behave normally on clean images, yet can be activated and output incorrect prediction if the input is stamped with a specific trigger pattern. Most existing backdoor attacks focus on manually defining imperceptible triggers in input space without considering the abnormality of triggers' latent representations in the poisoned model. These attacks are susceptible to backdoor detection algorithms and even visual inspection. In this paper, We propose a novel and stealthy backdoor attack - DEFEAT. It poisons the clean data using adaptive imperceptible perturbation and restricts latent representation during training process to strengthen our attack's stealthiness and resistance to defense algorithms. We conduct extensive experiments on multiple image classifiers using real-world datasets to demonstrate that our attack can 1) hold against the state-of-the-art defenses, 2) deceive the victim model with high attack success without jeopardizing model utility, and 3) provide practical stealthiness on image data.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Backdoor Attacks.
  • 2.2. Backdoor Defense.
  • 2.3. Threat Model
  • 3. Our Proposed Model: DEFEAT
  • 3.1. Preliminary
  • 3.2. Trigger Generation
  • 3.3. Backdoor Implantation
  • 4. Evaluation
  • 4.1. Attack Experiments
  • 4.2. Defense Experiments
  • 4.3. Hyperparameter Analysis
  • 5. Conclusion
  • 6. Acknowledgments
  • References

Knowls

  1. Knowl 1 — DEFEAT Bilevel Optimization Formulation

    model/method

    DEFEAT formulates the stealthy backdoor attack as a constrained bilevel optimization problem. For a deep neural network classifier Fθ:I→RKF_\theta: \mathcal{I} \to \mathbb{R}^K parameterized by θ\theta (where I⊂RC×H×W\mathcal{I} \subset \mathbb{R}^{C \times H \times W} is the image domain and KK is the number of classes), clean dataset Dc={(xi,yi)}i=1N\mathcal{D}_c = \{(x_i, y_i)\}_{i=1}^N, and poisoned dataset Db⊂Dc\mathcal{D}_b \subset \mathcal{D}_c, the objective is defined as:

    min⁡θβ1Lclean(θ)+β2Ladv(θ)\min_\theta \beta_1 \mathcal{L}_{clean}(\theta) + \beta_2 \mathcal{L}_{adv}(\theta)

    s.t.ϕ(θ)=arg⁡min⁡ϕ∑i=1N[Lce(Fθ(Tϕ(xi)),yt)+max⁡(d(Tϕ(xi),xi)−ϵ,0)]\text{s.t.} \quad \phi(\theta) = \arg\min_\phi \sum_{i=1}^N \left[ \mathcal{L}_{ce}(F_\theta(T_\phi(x_i)), y_t) + \max(d(T_\phi(x_i), x_i) - \epsilon, 0) \right]

    where β1\beta_1 and β2\beta_2 are balancing hyperparameter weights, Lce\mathcal{L}_{ce} is the standard cross-entropy loss, yty_t is the predefined target class label, d(Tϕ(x),x)=∥Tϕ(x)−x∥2d(T_\phi(x), x) = \|T_\phi(x) - x\|_2 is the ℓ2\ell_2-norm distance in pixel space, and ϵ>0\epsilon > 0 denotes the stealthiness perturbation budget.

    The clean data loss is:

    Lclean(θ)=∑(x,y)∈DcLce(Fθ(x),y)\mathcal{L}_{clean}(\theta) = \sum_{(x, y) \in \mathcal{D}_c} \mathcal{L}_{ce}(F_\theta(x), y)

    and the adversarial poisoning loss is:

    Ladv(θ)=∑(x,y)∈Db[Lce(Fθ(Tϕ(x)),yt)+Llf(x,Tϕ(x),Fθ)]\mathcal{L}_{adv}(\theta) = \sum_{(x, y) \in \mathcal{D}_b} \left[ \mathcal{L}_{ce}(F_\theta(T_\phi(x)), y_t) + \mathcal{L}_{lf}(x, T_\phi(x), F_\theta) \right]

    where Tϕ(x)=x+ϕT_\phi(x) = x + \phi is the backdoor injection function parameterized by the universal noise pattern ϕ\phi, and Llf\mathcal{L}_{lf} is the latent feature constraint loss.

  2. Knowl 2 — Latent Feature Constraint Loss via Intermediate Auxiliary Classifiers

    equation

    To prevent backdoor triggers from exhibiting abnormal, distinguishable activations in intermediate layers, DEFEAT trains an auxiliary linear classifier h(l)h^{(l)} on the intermediate representations of the clean model Fθ=F(N)∘⋯∘F(1)F_\theta = F^{(N)} \circ \dots \circ F^{(1)}. For intermediate layer l∈[1:N]l \in [1:N] with activation tensor z(l)∈RC(l)×H(l)×W(l)z^{(l)} \in \mathbb{R}^{C^{(l)} \times H^{(l)} \times W^{(l)}}, the auxiliary classifier is parameterized by weights ψ(l)∈RC(l)×K\psi^{(l)} \in \mathbb{R}^{C^{(l)} \times K} and bias ξ(l)∈RK\xi^{(l)} \in \mathbb{R}^K:

    h(l)(z(l))=pool(z(l))ψ(l)+ξ(l)h^{(l)}(z^{(l)}) = \text{pool}(z^{(l)})\psi^{(l)} + \xi^{(l)}

    where pool:RC(l)×H(l)×W(l)→RC(l)\text{pool}: \mathbb{R}^{C^{(l)} \times H^{(l)} \times W^{(l)}} \to \mathbb{R}^{C^{(l)}} is a global average pooling operator. During training of h(l)h^{(l)}, FθF_\theta acts as a frozen feature extractor trained on Dc\mathcal{D}_c via cross-entropy loss.

    The latent feature constraint loss Llfλ\mathcal{L}_{lf}^\lambda penalizes the discrepancy between clean and poisoned intermediate activations as evaluated by h(l)h^{(l)}:

    Llfλ=mean(∑l∈[1:N]λ(l)∣h(l)(zx(l))−h(l)(zTϕ(x)(l))∣)\mathcal{L}_{lf}^\lambda = \text{mean}\left( \sum_{l \in [1:N]} \lambda^{(l)} \left| h^{(l)}(z_x^{(l)}) - h^{(l)}(z_{T_\phi(x)}^{(l)}) \right| \right)

    where zx(l)z_x^{(l)} and zTϕ(x)(l)z_{T_\phi(x)}^{(l)} denote the features extracted at the ll-th layer for clean image xx and poisoned image Tϕ(x)T_\phi(x), respectively, and λ=(λ(1),…,λ(N))\lambda = (\lambda^{(1)}, \dots, \lambda^{(N)}) is a non-negative weighting vector satisfying ∑l=1Nλ(l)=1\sum_{l=1}^N \lambda^{(l)} = 1.

  3. Knowl 3 — DEFEAT Backdoor Attack Training Procedure

    algorithm

    The DEFEAT poisoning algorithm alternates between optimizing the universal noise trigger pattern ϕ\phi and updating the classifier parameters θ\theta under the latent feature constraint.

    Input: Clean Dataset Dc\mathcal{D}_c, Loss weights β1,β2\beta_1, \beta_2, Injection ratio η\eta, Total alternating steps RR, Stealthiness budget ϵ\epsilon
    Output: Backdoored classifier parameters θ\theta, Backdoor injection trigger ϕ\phi
    Train clean model FθF_\theta on Dc\mathcal{D}_c
    for each layer l∈[1:N]l \in [1 : N] do
        Train auxiliary intermediate classifier h(l)h^{(l)} on clean data Dc\mathcal{D}_c while keeping FθF_\theta fixed
    end for
    Initialize ϕ\phi
    Sample poisoning subset Db⊂Dc\mathcal{D}_b \subset \mathcal{D}_c such that ∣Db∣/∣Dc∣=η|\mathcal{D}_b| / |\mathcal{D}_c| = \eta
    for i=1i = 1 to RR do
        Optimize ϕ\phi via arg⁡min⁡ϕ∑(xj,yj)∈Dc[Lce(Fθ(xj+ϕ),yt)+max⁡(∥ϕ∥2−ϵ,0)]\arg\min_\phi \sum_{(x_j, y_j) \in \mathcal{D}_c} [\mathcal{L}_{ce}(F_\theta(x_j + \phi), y_t) + \max(\|\phi\|_2 - \epsilon, 0)]
        Compute total training loss L=β1Lclean(θ)+β2Ladv(θ)\mathcal{L} = \beta_1 \mathcal{L}_{clean}(\theta) + \beta_2 \mathcal{L}_{adv}(\theta)
        Update θ\theta via stochastic gradient descent
    end for
    return θ,ϕ\theta, \phi

    In standard implementations, R=20R = 20, β1=1\beta_1 = 1, β2=0.1\beta_2 = 0.1, the injection ratio is η=0.1\eta = 0.1, and model updates in line 8 use SGD with a fine-tuning learning rate of 0.0010.001.

  4. Knowl 4 — Attack Effectiveness across Benchmark Datasets and Architectures

    data/table

    The attack effectiveness of DEFEAT was evaluated against BadNets, SIG, ReFool, WaNet, and DFST across CIFAR-10, GTSRB, and a 10-class subset of ImageNet using VGG16, ResNet34, and WideResNet architectures. Performance is measured via Test Accuracy Rate (TAR, %) on clean samples and Attack Success Rate (ASR, %) on poisoned samples (injection ratio η=0.1\eta = 0.1, target label yt=0y_t = 0, averaged over 10 independent runs):

    Model Attack CIFAR-10 GTSRB ImageNet
    TAR ASR TAR ASR TAR ASR
    Clean Model 92.91±0.3092.91\pm0.30 - 98.43±0.0798.43\pm0.07 - 85.26±0.3085.26\pm0.30 -
    BadNets 90.03±0.4390.03\pm0.43 99.17±0.0299.17\pm0.02 95.43±0.0695.43\pm0.06 100.00±0.00100.00\pm0.00 83.54±0.2083.54\pm0.20 97.69±0.1097.69\pm0.10
    SIG 87.33±0.3087.33\pm0.30 99.75±0.0199.75\pm0.01 95.81±0.0595.81\pm0.05 99.85±0.0199.85\pm0.01 82.25±0.1282.25\pm0.12 98.92±0.1298.92\pm0.12
    VGG16 ReFool 87.24±0.7087.24\pm0.70 99.58±0.0399.58\pm0.03 94.66±0.0594.66\pm0.05 94.15±0.0194.15\pm0.01 77.90±0.0277.90\pm0.02 98.33±0.0298.33\pm0.02
    WaNet 91.87±0.5591.87\pm0.55 99.95±0.0299.95\pm0.02 96.59±0.1596.59\pm0.15 96.41±0.0296.41\pm0.02 83.30±0.3083.30\pm0.30 98.81±0.0298.81\pm0.02
    DFST 90.17±0.2590.17\pm0.25 99.80±0.0299.80\pm0.02 96.78±0.2596.78\pm0.25 98.96±0.0198.96\pm0.01 81.38±0.5081.38\pm0.50 99.96±0.0299.96\pm0.02
    DEFEAT (Ours) 91.93±0.5091.93\pm0.50 99.30±0.0199.30\pm0.01 96.24±0.0596.24\pm0.05 98.76±0.0298.76\pm0.02 84.97±0.1884.97\pm0.18 99.94±0.0199.94\pm0.01
    Clean Model 92.33±0.0192.33\pm0.01 - 98.42±0.0698.42\pm0.06 - 84.66±0.2584.66\pm0.25 -
    BadNets 91.73±0.0391.73\pm0.03 99.86±0.0199.86\pm0.01 97.90±0.4097.90\pm0.40 98.91±0.1198.91\pm0.11 79.45±0.1679.45\pm0.16 94.33±0.0194.33\pm0.01
    SIG 91.49±0.0391.49\pm0.03 99.38±0.2099.38\pm0.20 98.19±0.0698.19\pm0.06 100.00±0.00100.00\pm0.00 80.37±0.3080.37\pm0.30 97.04±0.0197.04\pm0.01
    ResNet34 ReFool 91.09±0.2291.09\pm0.22 99.05±0.1099.05\pm0.10 97.94±0.0197.94\pm0.01 98.47±0.0398.47\pm0.03 74.39±0.2474.39\pm0.24 94.35±0.2294.35\pm0.22
    WaNet 92.03±0.3092.03\pm0.30 99.96±0.0199.96\pm0.01 98.19±0.3198.19\pm0.31 99.83±0.0199.83\pm0.01 79.77±0.3079.77\pm0.30 95.31±0.0395.31\pm0.03
    DFST 91.21±0.4391.21\pm0.43 99.85±0.0399.85\pm0.03 97.39±0.3397.39\pm0.33 98.99±0.0198.99\pm0.01 80.74±0.3780.74\pm0.37 98.59±0.0398.59\pm0.03
    DEFEAT (Ours) 92.25±0.2592.25\pm0.25 99.98±0.0299.98\pm0.02 98.26±0.3098.26\pm0.30 99.01±0.0599.01\pm0.05 82.63±0.2282.63\pm0.22 98.98±0.0198.98\pm0.01
    Clean Model 93.35±1.1093.35\pm1.10 - 98.42±0.0698.42\pm0.06 - 85.79±0.1085.79\pm0.10 -
    BadNets 92.54±0.0392.54\pm0.03 99.88±0.0299.88\pm0.02 97.39±0.0297.39\pm0.02 99.98±0.0199.98\pm0.01 84.02±0.1384.02\pm0.13 91.61±0.0691.61\pm0.06
    SIG 91.73±0.0291.73\pm0.02 99.90±0.0199.90\pm0.01 97.74±0.0297.74\pm0.02 96.99±0.0096.99\pm0.00 82.84±0.2282.84\pm0.22 98.22±0.0198.22\pm0.01
    WideResNet ReFool 90.65±0.2090.65\pm0.20 99.80±0.3099.80\pm0.30 97.53±0.1397.53\pm0.13 96.62±0.3396.62\pm0.33 82.57±0.7582.57\pm0.75 95.37±0.0195.37\pm0.01
    WaNet 91.63±0.5091.63\pm0.50 99.57±0.3099.57\pm0.30 97.54±0.2197.54\pm0.21 97.56±0.0197.56\pm0.01 83.69±0.6083.69\pm0.60 95.42±0.1095.42\pm0.10
    DFST 92.14±0.4092.14\pm0.40 99.85±0.0299.85\pm0.02 97.04±0.3397.04\pm0.33 98.19±0.0298.19\pm0.02 83.92±0.3383.92\pm0.33 99.89±0.0299.89\pm0.02
    DEFEAT (Ours) 93.24±0.7093.24\pm0.70 99.98±0.1199.98\pm0.11 98.55±0.0198.55\pm0.01 99.90±0.0199.90\pm0.01 84.08±0.3084.08\pm0.30 99.88±0.1499.88\pm0.14

    DEFEAT consistently maintains clean accuracy (TAR) close to the clean unpoisoned model baseline (e.g., 93.24%93.24\% vs 93.35%93.35\% on CIFAR-10 with WideResNet) while achieving attack success rates (ASR) consistently above 98.76%98.76\% across all settings.

  5. Knowl 5 — Quantitative Attack Stealthiness Metrics

    data/table

    Image-level and perceptual stealthiness of poisoned images generated by DEFEAT and baseline attacks were evaluated on 500 randomly selected test samples across CIFAR-10, GTSRB, and ImageNet using Structural Similarity Index (SSIM ↑\uparrow), Peak Signal-to-Noise Ratio (PSNR ↑\uparrow, dB), and Learned Perceptual Image Patch Similarity (LPIPS ↓\downarrow, using pre-trained AlexNet features):

    Dataset Metric BadNets ReFool SIG WaNet DFST DEFEAT (Ours)
    SSIM ↑\uparrow 0.97630.9763 0.65420.6542 0.95780.9578 0.88540.8854 0.72100.7210 0.9813\mathbf{0.9813}
    CIFAR-10 LPIPS ↓\downarrow 0.00120.0012 0.06970.0697 0.00150.0015 0.00900.0090 0.08810.0881 0.0009\mathbf{0.0009}
    PSNR ↑\uparrow 20.0620.06 18.3718.37 25.2325.23 19.3019.30 16.7916.79 21.3121.31
    SSIM ↑\uparrow 0.95010.9501 0.74180.7418 0.71030.7103 0.9669\mathbf{0.9669} 0.75940.7594 0.91710.9171
    GTSRB LPIPS ↓\downarrow 0.0303\mathbf{0.0303} 0.30970.3097 0.08500.0850 0.05840.0584 0.25100.2510 0.04270.0427
    PSNR ↑\uparrow 23.4123.41 20.5720.57 25.2825.28 30.1130.11 21.5821.58 30.25\mathbf{30.25}
    SSIM ↑\uparrow 0.9955\mathbf{0.9955} 0.85640.8564 0.86800.8680 0.93590.9359 0.71290.7129 0.97650.9765
    ImageNet LPIPS ↓\downarrow 0.0062\mathbf{0.0062} 0.45740.4574 0.05730.0573 0.03600.0360 0.21050.2105 0.01590.0159
    PSNR ↑\uparrow 32.9532.95 20.4220.42 25.3025.30 29.5929.59 23.0123.01 37.50\mathbf{37.50}

    DEFEAT yields superior perceptual similarity on CIFAR-10 (lowest LPIPS 0.00090.0009 and highest SSIM 0.98130.9813) and achieves the highest PSNR across GTSRB (30.2530.25) and ImageNet (37.5037.50). While BadNets achieves competitive SSIM/LPIPS scores on larger images due to its localized small square trigger, it creates localized visual artifacts, whereas DEFEAT's Grad-CAM activations align directly with the main semantic object of the clean image.

  6. Knowl 6 — Robustness against Neural Attention Distillation Defense

    data/table

    Neural Attention Distillation (NAD) utilizes a clean teacher model to fine-tune a backdoored student network to align intermediate attention maps and erase backdoor triggers. The defense was evaluated over 20 distillation epochs on ImageNet using a ResNet34 classifier:

    Attack Original Epoch #5 Epoch #10 Epoch #15 Epoch #20
    TAR ASR TAR ASR TAR ASR TAR ASR TAR ASR
    BadNets 79.4579.45 94.3394.33 62.7862.78 14.7914.79 75.3175.31 14.4214.42 75.5275.52 13.9913.99 74.7274.72 7.487.48
    ReFool 80.3780.37 97.0497.04 67.8967.89 34.4834.48 76.6576.65 18.2418.24 76.4976.49 14.4714.47 76.4476.44 14.5214.52
    SIG 74.3974.39 94.3594.35 73.8673.86 45.4545.45 74.1974.19 35.5035.50 74.7774.77 31.8731.87 74.9374.93 28.2828.28
    WaNet 79.7779.77 95.3195.31 66.9366.93 36.4236.42 57.2857.28 35.1335.13 69.5769.57 32.9932.99 69.7169.71 36.0136.01
    DFST 80.7480.74 98.5998.59 77.9777.97 50.7850.78 76.9076.90 42.0142.01 79.1379.13 36.0636.06 79.8579.85 31.7131.71
    DEFEAT (Ours) 82.6382.63 98.9898.98 80.9080.90 86.7186.71 82.4682.46 81.5381.53 77.8977.89 82.2182.21 79.5379.53 79.9979.99

    While baseline attacks suffer drastic ASR reductions under NAD (e.g., BadNets drops from 94.33%94.33\% to 7.48%7.48\% and DFST drops from 98.59%98.59\% to 31.71%31.71\% at epoch 20), DEFEAT maintains a high ASR of 79.99%79.99\% and clean TAR of 79.53%79.53\%, resisting distillation-based trigger erasure.

  7. Knowl 7 — Resistance to STRIP, Neural Cleanse, and Fine-Pruning Defenses

    empirical result

    DEFEAT was tested against three representative backdoor defense mechanisms:

    1. STRIP (STRong Intentional Perturbation): STRIP detects poisoned inputs by computing the entropy of output probabilities when perturbed by superimposing random benign inputs. On ResNet34 evaluated on CIFAR-10 and GTSRB, the normalized entropy probability distributions of DEFEAT clean and poisoned samples overlap almost completely, rendering poisoned inputs indistinguishable from clean inputs (unlike BadNets, which exhibits separated entropy distributions).

    2. Neural Cleanse: Neural Cleanse detects backdoors via trigger reverse engineering and flags models with an anomaly index ≥2.0\ge 2.0. On ResNet34 across CIFAR-10, GTSRB, and ImageNet, DEFEAT maintains anomaly indices below the 2.02.0 detection threshold, whereas BadNets (3.843.84) and SIG (2.272.27) are flagged as anomalies on ImageNet. Furthermore, triggers reverse-engineered from DEFEAT models show no structured pattern and resemble clean representations.

    3. Fine-Pruning: Neurons in the final convolutional layer of a VGG16 model were pruned based on activation values. When the pruning rate reached 95%95\%, DEFEAT preserved an ASR above 90%90\%, while clean TAR dropped sharply to approximately 60%60\%, indicating that the backdoor cannot be pruned away without destroying benign utility.

  8. Knowl 8 — Hyperparameter Sensitivity: Injection Ratio and Stealthiness Budget

    data/table

    The effects of injection ratio η\eta and stealthiness perturbation budget ϵ\epsilon on Test Accuracy Rate (TAR, %) and Attack Success Rate (ASR, %) were analyzed using a ResNet34 classifier:

    Injection Ratio η\eta on GTSRB:

    Injection Ratio (η\eta) TAR (%) ASR (%)
    1%1\% 98.79±0.0798.79 \pm 0.07 93.66±0.2493.66 \pm 0.24
    5%5\% 98.56±0.0598.56 \pm 0.05 96.72±0.1496.72 \pm 0.14
    10%10\% 98.26±0.3098.26 \pm 0.30 99.01±0.0599.01 \pm 0.05
    15%15\% 97.53±0.0397.53 \pm 0.03 99.21±0.0299.21 \pm 0.02

    Perturbation Budget ϵ\epsilon across Datasets:

    Dataset Budget (ϵ\epsilon) TAR (%) ASR (%)
    2.02.0 92.65±0.2292.65 \pm 0.22 99.63±0.0199.63 \pm 0.01
    CIFAR-10 1.01.0 92.25±0.2592.25 \pm 0.25 99.98±0.0299.98 \pm 0.02
    0.50.5 91.05±0.2091.05 \pm 0.20 97.93±0.0297.93 \pm 0.02
    3.03.0 98.41±0.1198.41 \pm 0.11 99.61±0.0199.61 \pm 0.01
    GTSRB 2.02.0 98.26±0.3098.26 \pm 0.30 99.01±0.0599.01 \pm 0.05
    1.01.0 98.13±0.4798.13 \pm 0.47 98.59±0.0698.59 \pm 0.06
    5.05.0 84.29±0.4084.29 \pm 0.40 99.91±0.0399.91 \pm 0.03
    ImageNet 3.03.0 82.63±0.2282.63 \pm 0.22 98.98±0.0198.98 \pm 0.01
    1.01.0 81.64±0.3381.64 \pm 0.33 99.62±0.0199.62 \pm 0.01

    Increasing η\eta elevates ASR above 99%99\% at the cost of a minor reduction in clean TAR. Increasing ϵ\epsilon eases backdoor implantation but reduces imperceptibility. Optimal stealth-effectiveness trade-offs are obtained at default budgets of ϵ=1.0\epsilon = 1.0 for CIFAR-10, ϵ=2.0\epsilon = 2.0 for GTSRB, and ϵ=3.0\epsilon = 3.0 for ImageNet.

  9. Knowl 9 — Experimental Setup and Architecture Configurations

    experimental setup

    The experimental evaluation used the following settings:

    • Datasets: CIFAR-10 (60k60\text{k} 32×3232 \times 32 images, 1010 classes, 50k50\text{k} train / 10k10\text{k} test); GTSRB (51.2k51.2\text{k} 40×4040 \times 40 traffic sign images, 4343 classes); ImageNet (1010 randomly chosen classes, 10.5k10.5\text{k} 224×224224 \times 224 train / 1.8k1.8\text{k} test images). All pixel values are normalized to [0,1][0, 1].
    • Target Class: Single-target attack targeting class index yt=0y_t = 0.
    • Victim Networks: VGG16, ResNet34, and WideResNet.
    • Training Hyperparameters: Clean base models are trained for 200200 epochs using SGD with initial learning rate 0.010.01, decayed by a factor of 0.10.1 at epochs 50,100,15050, 100, 150. To constrain poisoning, the last convolutional layer is selected for VGG16, the last three residual blocks for ResNet34 (equal weights λ(l)=1/3\lambda^{(l)} = 1/3), and the last five residual blocks for WideResNet (equal weights λ(l)=1/5\lambda^{(l)} = 1/5). Fine-tuning during poisoning uses a learning rate of 0.0010.001, β1=1\beta_1 = 1, and β2=0.1\beta_2 = 0.1.

Coverage note — None was omitted; all contributed models, optimization objectives, algorithms, benchmark tables, defense evaluations, and hyperparameter analyses are fully covered.

References

  1. 1.Eugene Bagdasaryan, Andreas Veit, Yiqing Hua, Deborah Estrin, and Vitaly Shmatikov. How to backdoor federated learning. In International Conference on Artificial Intelligence and Statistics, pages 2938–2948. PMLR, 2020.
  2. 2.Mauro Barni, Kassem Kallas, and Benedetta Tondi. A new backdoor attack in cnns by training set corruption without label poisoning. In 2019 IEEE International Conference on Image Processing (ICIP), pages 101–105. IEEE, 2019.
  3. 3.Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Benjamin Edwards, Taesung Lee, Ian Molloy, and Biplav Srivastava. Detecting backdoor attacks on deep neural networks by activation clustering. arXiv preprint arXiv:1811.03728, 2018.
  4. 4.Huili Chen, Cheng Fu, Jishen Zhao, and Farinaz Koushanfar. Deepinspect: A black-box trojan detection and mitigation framework for deep neural networks. In IJCAI, pages 4658–4664, 2019.
  5. 5.Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. Targeted backdoor attacks on deep learning systems using data poisoning. arXiv preprint arXiv:1712.05526, 2017.
  6. 6.Siyuan Cheng, Yingqi Liu, Shiqing Ma, and Xiangyu Zhang. Deep feature space trojan attack of neural networks by controlled detoxification. arXiv preprint arXiv:2012.11212, 2020.
  7. 7.Khoa Doan, Yingjie Lao, Weijie Zhao, and Ping Li. Lira: Learnable, imperceptible and robust backdoor attacks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 11966–11976, 2021.
  8. 8.Yansong Gao, Change Xu, Derui Wang, Shiping Chen, Damith C Ranasinghe, and Surya Nepal. Strip: A defence against trojan attacks on deep neural networks. In Proceedings of the 35th Annual Computer Security Applications Conference, pages 113–125, 2019.
  9. 9.Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. Badnets: Identifying vulnerabilities in the machine learning model supply chain. arXiv preprint arXiv:1708.06733, 2017.
  10. 10.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In European conference on computer vision, pages 630–645. Springer, 2016.
  12. 12.Quan Huynh-Thu and Mohammed Ghanbari. Scope of validity of psnr in image/video quality assessment. Electronics letters, 44(13):800–801, 2008.
  13. 13.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  14. 14.Shaofeng Li, Minhui Xue, Benjamin Zhao, Haojin Zhu, and Xinpeng Zhang. Invisible backdoor attacks on deep neural networks via steganography and regularization. IEEE Transactions on Dependable and Secure Computing, 2020.
  15. 15.Yuezun Li, Yiming Li, Baoyuan Wu, Longkang Li, Ran He, and Siwei Lyu. Invisible backdoor attack with samplespecific triggers. In IEEE International Conference on Computer Vision (ICCV), 2021.
  16. 16.Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. Neural attention distillation: Erasing backdoor triggers from deep neural networks. In ICLR, 2021.
  17. 17.Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. Finepruning: Defending against backdooring attacks on deep neural networks. In International Symposium on Research in Attacks, Intrusions, and Defenses, pages 273–294. Springer, 2018.
  18. 18.Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang. Trojaning attack on neural networks. 2017.
  19. 19.Yunfei Liu, Xingjun Ma, James Bailey, and Feng Lu. Reflection backdoor: A natural backdoor attack on deep neural networks. In European Conference on Computer Vision, pages 182–199. Springer, 2020.
  20. 20.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083, 2017.
  21. 21.Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  22. 22.Anh Nguyen and Anh Tran. Input-aware dynamic backdoor attack. arXiv preprint arXiv:2010.08138, 2020.
  23. 23.Anh Nguyen and Anh Tran. Wanet – imperceptible warpingbased backdoor attack. 2021.
  24. 24.Ximing Qiao, Yukun Yang, and Hai Li. Defending neural backdoors via generative distribution modeling. arXiv preprint arXiv:1910.04749, 2019.
  25. 25.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  26. 26.Aniruddha Saha, Akshayvarun Subramanya, and Hamed Pirsiavash. Hidden trigger backdoor attacks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11957–11965, 2020.
  27. 27.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, pages 618–626, 2017.
  28. 28.Ali Shafahi, W Ronny Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, and Tom Goldstein. Poison frogs! targeted clean-label poisoning attacks on neural networks. arXiv preprint arXiv:1804.00792, 2018.
  29. 29.Reza Shokri et al. Bypassing backdoor detection algorithms in deep learning. In 2020 IEEE European Symposium on Security and Privacy (EuroS&P), pages 175–183. IEEE, 2020.
  30. 30.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  31. 31.Johannes Stallkamp, Marc Schlipsing, Jan Salmen, and Christian Igel. Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition. Neural networks, 32:323–332, 2012.
  32. 32.Jacob Steinhardt, Pang Wei Koh, and Percy Liang. Certified defenses for data poisoning attacks. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pages 3520–3532, 2017.
  33. 33.Ziteng Sun, Peter Kairouz, Ananda Theertha Suresh, and H Brendan McMahan. Can you really backdoor federated learning? arXiv preprint arXiv:1911.07963, 2019.
  34. 34.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks. In Advances in neural information processing systems, pages 3104–3112, 2014.
  35. 35.Brandon Tran, Jerry Li, and Aleksander Madry. Spectral signatures in backdoor attacks. arXiv preprint arXiv:1811.00636, 2018.
  36. 36.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  37. 37.Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao. Neural cleanse: Identifying and mitigating backdoor attacks in neural networks. In 2019 IEEE Symposium on Security and Privacy (SP), pages 707–723. IEEE, 2019.
  38. 38.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004.
  39. 39.Chulin Xie, Keli Huang, Pin-Yu Chen, and Bo Li. Dba: Distributed backdoor attacks against federated learning. In International Conference on Learning Representations, 2019.
  40. 40.Yuanshun Yao, Huiying Li, Haitao Zheng, and Ben Y Zhao. Latent backdoor attacks on deep neural networks. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security, pages 2041–2055, 2019.
  41. 41.Yunrui Yu, Xitong Gao, and Cheng-Zhong Xu. Lafeat: Piercing through adversarial defenses with latent features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5735–5745, 2021.
  42. 42.Sergey Zagoruyko and Nikos Komodakis. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016.
  43. 43.Yi Zeng, Won Park, Z Morley Mao, and Ruoxi Jia. Rethinking the backdoor attacks’ triggers: A frequency perspective. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16473–16481, 2021.
  44. 44.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018.
  45. 45.Shihao Zhao, Xingjun Ma, Xiang Zheng, James Bailey, Jingjing Chen, and Yu-Gang Jiang. Clean-label backdoor attacks on video recognition models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14443–14452, 2020.

Citation

MLA
Zhao, Z., et al. “DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 15192–201, https://doi.org/10.1109/CVPR52688.2022.01478.
APA
Zhao, Z., Chen, X., Xuan, Y., Dong, Y., Wang, D., & Liang, K. (2022). DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15192–15201. https://doi.org/10.1109/CVPR52688.2022.01478
Chicago
Zhao, Z., X. Chen, Y. Xuan, Y. Dong, D. Wang, and K. Liang. 2022. “DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15192–201. https://doi.org/10.1109/CVPR52688.2022.01478.
Harvard
Zhao, Z. et al. (2022) “DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 15192–15201. Available at: https://doi.org/10.1109/CVPR52688.2022.01478.
Vancouver
1. Zhao Z, Chen X, Xuan Y, Dong Y, Wang D, Liang K (2022) DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 15192–15201

BibTeX

@inproceedings{Zhao_2022, title={DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints}, url={http://dx.doi.org/10.1109/CVPR52688.2022.01478}, DOI={10.1109/cvpr52688.2022.01478}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhao, Zhendong and Chen, Xiaojun and Xuan, Yuexin and Dong, Ye and Wang, Dakui and Liang, Kaitai}, year={2022}, month=June, pages={15192–15201} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE