Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains

Qilong ZhangXiaodan LiYuefeng ChenJingkuan SongLianli GaoYuan HeHui Xue

article2022ICLR88 citations

Develops a generative framework that disrupts low-level image features using only ImageNet knowledge, enabling transferable black-box attacks against models deployed on completely unknown target domains.

Listen

Deep neural networks are widely used in mission-critical vision applications, yet they remain vulnerable to adversarial attacks—subtle, engineered perturbations to input images that cause models to make incorrect predictions. Previous security assessments primarily assumed that attackers have access to the target model's training data or can repeatedly probe the deployed system with queries. However, real-world systems are rarely accessible for frequent querying, and their underlying training datasets are usually kept private. This creates an urgent operational need to evaluate whether vision systems can be compromised under strict black-box conditions, where attackers have zero knowledge of the target data, model architecture, or task.

The main objective of the article is to demonstrate that attackers can craft highly transferable adversarial examples for unknown black-box domains using only a substitute model trained on standard, publicly available ImageNet data. It introduces the Beyond ImageNet Attack framework and evaluates its ability to fool diverse, unknown target classifiers across eight distinct image classification tasks.

To conduct this evaluation, the researchers trained a generative model on ImageNet to learn adversarial perturbations that disrupt general, low-level visual features in intermediate network layers rather than task-specific final layers. The study tested the approach across four substitute architectures and evaluated attack success against seven target domains, comprising four coarse-grained datasets and three fine-grained datasets. To bridge the gap between source and target domains, the authors introduced two enhancements: a random normalization module to simulate varying data distributions and a domain-agnostic attention module to focus perturbations on core object features.

The analysis produced several key findings. First, the proposed method substantially outperforms existing state-of-the-art transfer attacks across all tested domains. In fine-grained classification tasks, the domain-agnostic attention variant reduced target model accuracy to an average of 36.42%, outperforming the strongest baseline by roughly 25.9 percentage points. In coarse-grained tasks, the random normalization variant reduced target accuracy to an average of 57.30%, beating baseline methods by 7.71 percentage points. Second, attacking intermediate and shallow network layers proved significantly more effective for cross-domain transfer than attacking deep, task-specific layers. Third, combining substitute models into an ensemble further increased attack strength, while standard data augmentation failed to provide similar transfer benefits.

These findings carry critical security and risk implications. They demonstrate that withholding training data, hiding model architectures, and restricting query access provide a false sense of security. Commercially deployed vision systems can be fooled in real time with a single forward pass by adversaries leveraging public datasets. Consequently, security assessments that assume isolated models are safe from transfer attacks significantly underestimate operational vulnerabilities in sensitive environments such as autonomous navigation, healthcare diagnostics, and automated surveillance.

Based on these results, machine learning practitioners and security teams should immediately incorporate cross-domain adversarial testing into their model validation pipelines rather than relying solely on standard robustness evaluations. Defenses must focus on hardening low-level feature representations and implementing robust input pre-processing. Combining random normalization and domain-agnostic attention should be tailored cautiously, as the article observed that using both simultaneously can produce mixed results depending on the underlying network backbone.

The primary limitation of the study is its focus on image classification benchmarks using standard convolutional backbones, leaving potential impacts on other vision tasks—such as object detection and segmentation—untested. Furthermore, generating effective cross-domain attacks requires training on a sufficiently large and diverse source dataset like ImageNet; models trained on smaller datasets transfer poorly. Nevertheless, given the consistent outperformance across multiple benchmarks and repeated runs with low standard deviations, confidence in the demonstrated vulnerabilities remains very high.

Cover for Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains

Abstract

Adversarial examples have posed a severe threat to deep neural networks due to their transferable nature. Currently, various works have paid great efforts to enhance the cross-model transferability, which mostly assume the substitute model is trained in the same domain as the target model. However, in reality, the relevant information of the deployed model is unlikely to leak. Hence, it is vital to build a more practical black-box threat model to overcome this limitation and evaluate the vulnerability of deployed models. In this paper, with only the knowledge of the ImageNet domain, we propose a Beyond ImageNet Attack (BIA) to investigate the transferability towards black-box domains (unknown classification tasks). Specifically, we leverage a generative model to learn the adversarial function for disrupting low-level features of input images. Based on this framework, we further propose two variants to narrow the gap between the source and target domains from the data and model perspectives, respectively. Extensive experiments on coarse-grained and fine-grained domains demonstrate the effectiveness of our proposed methods. Notably, our methods outperform state-of-the-art approaches by up to 7.71% (towards coarse-grained domains) and 25.91% (towards fine-grained domains) on average. Our code is available at \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Transferable Adversarial Examples beyond ImageNet
  • 3.1 Problem Formulation
  • 3.2 Preliminary
  • 3.3 Random Normalization Module
  • 3.4 Domain-agnostic Attention Module
  • 4 Experiments
  • 4.1 Transferability Comparisons
  • 4.1.1 Results on Coarse-grained Domain
  • 4.1.2 Results on Fine-grained Domain
  • 4.1.3 Results on Source Domain
  • 4.2 Combination of Domain-agnostic Attention and Random Normalization
  • 5 Conclusion
  • 6 Acknowledge
  • References
  • A Appendix
  • A.1 Select Layer for Attacking
  • A.2 Select Gaussian Distribution for Random Normalization
  • A.3 Data Augmentation vs. Random Normalization
  • A.4 Insight into the Generator
  • A.5 Discussion on Changing Source Domain
  • A.6 Discussion on Ensemble-model Attacks
  • A.7 Discussion on Standard Deviation across Multiple Random Runs
  • A.8 Effects of ℛ​𝒩\mathcal{RN} and 𝒟​𝒜\mathcal{DA} on Coarse-grained and Fine-grained Tasks

Knowls

  1. Knowl 1 — Cross-Domain Black-Box Adversarial Attack Formulation

    definition

    In a cross-domain black-box adversarial attack setting, an attacker crafts human-imperceptible perturbations for images drawn from an unknown target domain distribution χt\chi_t to fool an inaccessible target classifier ft(⋅)f_t(\cdot), relying solely on the knowledge of a source domain data distribution χs\chi_s (such as ImageNet) and a pre-trained substitute model fs(⋅)f_s(\cdot) trained on χs\chi_s. The attacker has no knowledge of χt\chi_t or ft(⋅)f_t(\cdot), and query access to ft(⋅)f_t(\cdot) is strictly prohibited.

    Formally, given a threat model Mθ∗M_{\theta^*} whose parameter θ∗\theta^* is optimized exclusively on the source domain, and a clean target domain image xt∼χtx_t \sim \chi_t, the goal is to satisfy: ft(Mθ∗(xt))≠ft(xt)subject to∥Mθ∗(xt)−xt∥∞≤ϵf_t(M_{\theta^*}(x_t)) \neq f_t(x_t) \quad \text{subject to} \quad \|M_{\theta^*}(x_t) - x_t\|_\infty \le \epsilon where ϵ\epsilon denotes the maximum allowable ℓ∞\ell_\infty perturbation magnitude.

  2. Knowl 2 — Beyond ImageNet Attack (BIA) Architecture and Feature Disruption Objective

    model/method

    The Beyond ImageNet Attack (BIA) trains a feed-forward generator network GθG_\theta on the large-scale source domain χs\chi_s (ImageNet) to generate cross-domain transferable adversarial examples in a single forward pass. Rather than optimizing a domain-specific label classification loss (which causes overfitting to source categories), BIA disrupts low-level intermediate features by minimizing the cosine similarity between the feature representations of benign source images xs∈RN×Hs×Wsx_s \in \mathbb{R}^{N \times H_s \times W_s} and perturbed images xs′=Gθ(xs)x'_s = G_\theta(x_s) extracted at a chosen intermediate layer LL of a substitute model fsf_s: θ∗=arg⁡min⁡θLcos(fsL(xs′),fsL(xs))\theta^* = \arg\min_\theta \mathcal{L}_{cos}(f_s^L(x'_s), f_s^L(x_s)) where Lcos(u,v)=u⋅v∥u∥2∥v∥2\mathcal{L}_{cos}(u, v) = \frac{u \cdot v}{\|u\|_2 \|v\|_2} denotes the cosine similarity function.

    The generator GθG_\theta consists of downsampling convolutional blocks, residual blocks, and upsampling blocks. During inference on any target domain input image xt∈RN×Ht×Wtx_t \in \mathbb{R}^{N \times H_t \times W_t}, the adversarial example xt′x'_t is produced in a single step with clipping onto the ℓ∞\ell_\infty ϵ\epsilon-ball: xt′=min⁡(xt+ϵ,max⁡(Gθ∗(xt),xt−ϵ))x'_t = \min(x_t + \epsilon, \max(G_{\theta^*}(x_t), x_t - \epsilon))

  3. Knowl 3 — Random Normalization Module for Domain Distribution Simulation

    model/method

    Standard deep vision networks normalize input images using source domain statistics (for ImageNet, channel mean μ=[0.485,0.456,0.406]\mu = [0.485, 0.456, 0.406] and standard deviation σ=[0.229,0.224,0.225]\sigma = [0.229, 0.224, 0.225]). Because target domain image distributions χt\chi_t diverge significantly in mean and variance from χs\chi_s, generators trained with static source normalization fail to generalize across domains.

    The Random Normalization (RN\mathcal{RN}) module simulates diverse target data distributions during training by perturbing the normalization scales and offsets applied to source inputs: RN(xs)=σ⋅xs−μ′σ′+μ\mathcal{RN}(x_s) = \sigma \cdot \frac{x_s - \mu'}{\sigma'} + \mu where μ′\mu' and σ′\sigma' are random scaling vectors sampled per batch from normal distributions: μ′∼N(μmean′,μstd′),σ′∼N(σmean′,σstd′)\mu' \sim \mathcal{N}(\mu'_{mean}, \mu'_{std}), \quad \sigma' \sim \mathcal{N}(\sigma'_{mean}, \sigma'_{std}) with default hyperparameters set to μmean′=0.50\mu'_{mean} = 0.50, μstd′=0.08\mu'_{std} = 0.08, σmean′=0.75\sigma'_{mean} = 0.75, and σstd′=0.08\sigma'_{std} = 0.08. The generator training objective incorporating RN\mathcal{RN} becomes: θ∗=arg⁡min⁡θLcos(fsL(RN(xs′)),fsL(RN(xs)))\theta^* = \arg\min_\theta \mathcal{L}_{cos}(f_s^L(\mathcal{RN}(x'_s)), f_s^L(\mathcal{RN}(x_s)))

  4. Knowl 4 — Domain-Agnostic Attention Module for Feature Map Weighting

    model/method

    Intermediate feature maps at layer LL of a substitute model fsf_s trained on source data can contain channels biased towards source-specific textures rather than universal object geometry. To mitigate the influence of biased feature maps, the Domain-Agnostic Attention (DA\mathcal{DA}) module aggregates feature maps across all CC channels at layer LL via cross-channel average pooling: AL=∣∑i=0C−1[fsL(xs)]i∣CA^L = \frac{\left|\sum_{i=0}^{C-1} [f_s^L(x_s)]_i\right|}{C} where [fsL(xs)]i[f_s^L(x_s)]_i is the ii-th channel slice of the feature tensor fsL(xs)f_s^L(x_s), and ALA^L serves as a spatial attention map.

    During training, ALA^L weights the feature representations element-wise via the Hadamard product ⊙\odot, guiding the generator GθG_\theta to disrupt core object regions: θ∗=arg⁡min⁡θLcos(AL⊙fsL(xs′),AL⊙fsL(xs))\theta^* = \arg\min_\theta \mathcal{L}_{cos}(A^L \odot f_s^L(x'_s), A^L \odot f_s^L(x_s))

  5. Knowl 5 — Joint Training Objective Combining Random Normalization and Domain-Agnostic Attention

    equation

    When both the Random Normalization (RN\mathcal{RN}) module and the Domain-Agnostic Attention (DA\mathcal{DA}) module are integrated into the Beyond ImageNet Attack framework, the generator network parameters θ∗\theta^* are optimized via: θ∗=arg⁡min⁡θLcos(AL⊙fsL(RN(xs′)),AL⊙fsL(RN(xs)))\theta^* = \arg\min_\theta \mathcal{L}_{cos}(A^L \odot f_s^L(\mathcal{RN}(x'_s)), A^L \odot f_s^L(\mathcal{RN}(x_s))) where:

    • xs∈RN×Hs×Wsx_s \in \mathbb{R}^{N \times H_s \times W_s} is a clean image sampled from the ImageNet source distribution χs\chi_s, and xs′=Gθ(xs)x'_s = G_\theta(x_s) is the generated adversarial example.
    • RN(xs)=σ⋅xs−μ′σ′+μ\mathcal{RN}(x_s) = \sigma \cdot \frac{x_s - \mu'}{\sigma'} + \mu scales inputs with randomly sampled statistics μ′∼N(0.50,0.08)\mu' \sim \mathcal{N}(0.50, 0.08) and σ′∼N(0.75,0.08)\sigma' \sim \mathcal{N}(0.75, 0.08), using ImageNet default vectors μ\mu and σ\sigma.
    • fsL(⋅)f_s^L(\cdot) denotes the intermediate feature extractor at layer LL of the substitute model fsf_s.
    • AL=∣∑i=0C−1[fsL(xs)]i∣CA^L = \frac{\left|\sum_{i=0}^{C-1} [f_s^L(x_s)]_i\right|}{C} is the cross-channel average pooled attention map computed from fsL(xs)f_s^L(x_s), where CC is the channel dimension.
    • ⊙\odot represents the Hadamard product.
    • Lcos(u,v)=u⋅v∥u∥2∥v∥2\mathcal{L}_{cos}(u, v) = \frac{u \cdot v}{\|u\|_2 \|v\|_2} is the cosine similarity loss.
  6. Knowl 6 — Cross-Domain Adversarial Transferability on Coarse-Grained Datasets

    data/table

    Evaluating attack transferability from an ImageNet substitute model (L∞≤10L_\infty \le 10) across four coarse-grained target domains (CIFAR-10, CIFAR-100, STL-10, SVHN) demonstrates that Beyond ImageNet Attack (BIA) and its variants outperform iterative attacks (PGD, DIM, DR, SSP) and the generative Cross-Domain Attack (CDA).

    Substitute Model Attack CIFAR-10 CIFAR-100 STL-10 SVHN Average Top-1 Acc (%)
    - Clean 93.78 74.27 77.59 96.03 85.42
    VGG-16 PGD 79.63 48.02 74.32 94.66 74.16
    DIM 77.16 44.75 72.74 91.53 71.55
    DR 72.49 39.04 72.56 93.27 69.34
    SSP 68.54 33.63 72.77 93.98 67.23
    CDA 66.41 32.37 72.91 92.17 65.97
    BIA (Ours) 57.38 22.47 69.45 90.44 59.94
    BIA+DA (Ours) 55.16 21.71 70.00 91.76 59.66
    BIA+RN (Ours) 52.81 20.82 67.55 88.03 57.30
    Res-152 PGD 86.17 56.38 74.51 93.94 77.75
    DIM 80.50 48.03 71.20 90.87 72.65
    DR 78.86 48.62 71.66 93.26 73.10
    SSP 75.54 42.38 72.66 92.63 70.80
    CDA 66.47 39.30 69.81 88.09 65.92
    BIA (Ours) 65.49 33.48 69.91 89.46 64.59
    BIA+DA (Ours) 65.34 32.68 69.65 91.38 64.76
    BIA+RN (Ours) 61.23 32.84 68.04 85.79 61.98
    Dense-169 PGD 84.55 54.29 74.55 93.83 76.81
    DIM 80.89 49.06 72.64 89.47 73.02
    DR 78.24 48.67 70.75 93.20 72.72
    SSP 77.13 42.18 72.53 91.64 70.87
    CDA 67.75 35.03 69.00 88.76 65.14
    BIA (Ours) 72.02 38.99 69.80 86.12 66.73
    BIA+DA (Ours) 71.69 38.95 70.60 88.02 67.32
    BIA+RN (Ours) 66.67 34.41 68.79 81.54 62.85

    Top-1 classification accuracy after attack is reported (lower values indicate higher transferability). Across all coarse-grained targets, the RN\mathcal{RN} variant delivers the strongest transferability gains (reducing average accuracy to 57.30% with VGG-16 and 61.98% with Res-152, outperforming CDA by 8.67% and 3.94%, respectively), showing that random normalization compensates for domain shifts in low-resolution and differently normalized coarse-grained datasets.

  7. Knowl 7 — Cross-Domain Adversarial Transferability on Fine-Grained Datasets

    data/table

    When evaluating cross-domain attacks on fine-grained image recognition datasets (CUB-200-2011, Stanford Cars, FGVC Aircraft) using Destruction and Construction Learning (DCL) models with ResNet50, SENet154, and SE-ResNet101 backbones under ℓ∞≤10\ell_\infty \le 10, BIA and its DA\mathcal{DA} and RN\mathcal{RN} variants significantly outperform existing transfer attacks.

    Substitute Model Attack CUB-200-2011 Avg Stanford Cars Avg FGVC Aircraft Avg Overall Average Top-1 Acc (%)
    - Clean 86.91 93.56 92.07 90.85
    VGG-16 PGD 80.31 88.93 83.65 84.30
    DIM 67.82 79.05 67.60 71.49
    DR 81.88 90.84 86.02 86.25
    SSP 64.74 72.25 62.48 66.49
    CDA 67.73 77.68 64.42 69.94
    BIA (Ours) 47.92 59.89 45.38 51.07
    BIA+DA (Ours) 39.50 47.75 34.70 40.65
    BIA+RN (Ours) 42.45 48.88 38.17 43.17
    Res-152 PGD 75.07 86.37 77.98 79.81
    DIM 59.67 74.67 60.55 64.96
    DR 80.08 89.58 80.08 83.25
    SSP 60.57 71.94 67.02 66.51
    CDA 50.57 64.62 56.91 57.36
    BIA (Ours) 49.53 51.00 40.71 47.08
    BIA+DA (Ours) 39.45 51.29 34.07 41.60
    BIA+RN (Ours) 38.34 40.30 37.36 38.67
    Dense-169 PGD 80.06 89.20 82.93 84.06
    DIM 63.71 77.47 62.98 68.06
    DR 77.22 88.29 77.00 80.84
    SSP 50.49 56.57 38.88 48.65
    CDA 56.97 67.60 61.16 61.91
    BIA (Ours) 30.07 34.37 23.25 29.23
    BIA+DA (Ours) 25.82 25.74 10.69 20.75
    BIA+RN (Ours) 29.93 31.55 21.67 27.72

    Top-1 accuracy reflects post-attack classification accuracy (lower is better). For Dense-169, vanilla BIA decreases the overall average top-1 accuracy to 29.23% compared to 61.91% for CDA and 48.65% for SSP. Incorporating the Domain-Agnostic Attention (DA\mathcal{DA}) module reduces the average accuracy further to 20.75%, outperforming SSP by 27.90% on Dense-169 and by 25.91% across all models on average.

  8. Knowl 8 — In-Domain and Cross-Model Black-Box Transferability on ImageNet

    empirical result

    Although developed for cross-domain attack transferability, Beyond ImageNet Attack (BIA) and its variants also enhance white-box and cross-model black-box transferability within the source domain (ImageNet) under an ℓ∞≤10\ell_\infty \le 10 perturbation budget.

    When training generators against VGG-16:

    • Vanilla BIA attains an average top-1 accuracy of 24.86% across 7 evaluated ImageNet architectures (VGG-16 white-box at 1.55%, Dense-169 at 32.35%, VGG-19 at 3.61%, Res-50 at 25.36%, Res-152 at 42.98%, Dense-121 at 26.97%, and Inc-v3 at 41.20%), compared to 32.01% for CDA and 27.68% for SSP.
    • BIA+DA reduces average top-1 accuracy to 19.60% (VGG-16: 1.04%, Inc-v3: 34.54%).
    • BIA+RN reduces average top-1 accuracy to 17.87% (Res-50: 16.52%, Inc-v3: 28.54%).

    When training generators against Dense-169:

    • Vanilla BIA yields an average top-1 accuracy of 12.05% across models (Dense-169 white-box: 6.45%, Inc-v3: 38.58%), compared to 12.39% for CDA (Inc-v3: 43.78%) and 14.19% for SSP.
    • BIA+DA and BIA+RN reduce average top-1 accuracy to 7.34% (Inc-v3: 26.51%) and 7.36% (Inc-v3: 14.24%), respectively. On the architecturally distinct Inception-v3 model, BIA+RN decreases target model accuracy from 43.78% (CDA) to 14.24% (a 29.54% reduction).
  9. Knowl 9 — Impact of Intermediate Layer Depth on Attack Transferability

    empirical result

    The depth of the intermediate layer LL chosen in the substitute model during generator training governs a trade-off between cross-domain and in-domain transferability:

    1. Shallow to middle layers provide superior cross-domain transferability towards black-box target domains. When using ResNet-152 as the substitute model, perturbing the shallow layer Conv2_3 is most effective for transfer to coarse-grained domains, whereas attacking the middle layer Conv3_8 provides the lowest average top-1 accuracy on fine-grained target domains.
    2. Deep layers provide higher cross-model transferability within the source domain (ImageNet). Attacking deep layer Conv5_3 in ResNet-152 achieves the strongest attack performance when transferring across ImageNet models, but transfers poorly to out-of-domain datasets.

    This indicates that lower-level representations capture general, domain-agnostic feature structures, whereas higher-level layers capture domain-specific semantic categories. The optimal layer choices across architectures are Maxpool.3 for VGG-16 and VGG-19, Conv3_8 for ResNet-152, and DenseBlock.2 for DenseNet-169.

  10. Knowl 10 — Conditional Incompatibility Between Random Normalization and Domain-Agnostic Attention

    limitation

    Combining the Random Normalization (RN\mathcal{RN}) module and the Domain-Agnostic Attention (DA\mathcal{DA}) module simultaneously does not consistently produce compounding benefits and can degrade attack transferability depending on the substitute architecture and target domain:

    1. Architecture-dependent feature response: On VGG-16, RN\mathcal{RN} reinforces discriminative object feature activations, allowing DA\mathcal{DA} and RN\mathcal{RN} to cooperate effectively (reducing coarse-grained post-attack average accuracy to 50.29% and fine-grained to 40.46%). On DenseNet-169, however, RN\mathcal{RN} suppresses the intermediate activation response of essential object features. Consequently, applying DA\mathcal{DA} in tandem with RN\mathcal{RN} on DenseNet-169 causes the generator to misdirect its attention away from salient regions, yielding worse transferability than using either module alone.
    2. Coarse-grained domains: Across all substitute models on coarse-grained tasks (CIFAR-10, CIFAR-100, STL-10, SVHN), applying RN\mathcal{RN} alone outperforms the combined DA+RN\mathcal{DA}+\mathcal{RN} setup (for example, on ResNet-152, BIA+RN yields 61.98% average top-1 accuracy versus 63.85% for BIA+DA+RN).
  11. Knowl 11 — Effect of Source Domain Dataset Scale on Cross-Domain Attack Transferability

    empirical result

    The cross-domain transferability of generator-based attacks depends heavily on the scale and visual diversity of the source training distribution. When a generator is trained on a small source dataset (CUB-200-2011 with 5,794 images) against a DCL ResNet-50 substitute model under ℓ∞≤10\ell_\infty \le 10:

    • Transferability to coarse-grained domains degrades substantially compared to training on ImageNet: target accuracy on CIFAR-10 remains high at 82.62% for BIA, 82.94% for BIA+DA, and 83.84% for BIA+RN (compared to 52.81% when trained on ImageNet with VGG-16).
    • On SVHN, target accuracy remains at 91.67% for BIA and 87.28% for BIA+DA.
    • On ImageNet, target accuracy is reduced to 44.15% for BIA, 37.15% for BIA+DA, and 39.43% for BIA+RN, outperforming CDA (50.99%).

    While BIA variants maintain superiority over CDA on small-scale training sets, training on a large-scale, diverse dataset like ImageNet is essential to achieve strong out-of-domain transferability across coarse-grained tasks.

Coverage note — Ablation experiments on hyperparameter grid searches for the Gaussian distribution in the RN module and standard deviation tables across random seeds were omitted as non-essential implementation details.

References

  1. 1.Shumeet Baluja and Ian Fischer. Adversarial transformation networks: Learning to generate adversarial examples. CoRR, abs/1703.09387, 2017.
  2. 2.Wieland Brendel, Jonas Rauber, and Matthias Bethge. Decision-based adversarial attacks: Reliable attacks against black-box machine learning models. In ICLR, 2018.
  3. 3.Nicholas Carlini and David A. Wagner. Towards evaluating the robustness of neural networks. In Symposium on Security and Privacy, 2017.
  4. 4.Rich Caruana, Alexandru Niculescu-Mizil, Geoff Crew, and Alex Ksikes. Ensemble selection from libraries of models. In ICML, 2004.
  5. 5.Jianbo Chen, Michael I. Jordan, and Martin J. Wainwright. Hopskipjumpattack: A query-efficient decision-based attack. In SP, 2020.
  6. 6.Yue Chen, Yalong Bai, Wei Zhang, and Tao Mei. Destruction and construction learning for finegrained image recognition. In CVPR, 2019.
  7. 7.Adam Coates, Andrew Y. Ng, and Honglak Lee. An analysis of single-layer networks in unsupervised feature learning. In Geoffrey J. Gordon, David B. Dunson, and Miroslav Dud0qk (eds.), AISTATS, 2011.
  8. 8.Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial attacks with momentum. In CVPR, 2018.
  9. 9.Ranjie Duan, Xiaofeng Mao, A. Kai Qin, Yuefeng Chen, Shaokai Ye, Yuan He, and Yun Yang. Adversarial laser beam: Effective physical-world attack to dnns in a blink. In CVPR, 2021.
  10. 10.Lianli Gao, Qilong Zhang, Jingkuan Song, Xianglong Liu, and Hengtao Shen. Patch-wise attack for fooling deep neural network. In ECCV, 2020a.
  11. 11.Lianli Gao, Qilong Zhang, Jingkuan Song, and Heng Tao Shen. Patch-wise++ perturbation for adversarial targeted attacks. CoRR, abs/2012.15503, 2020b.
  12. 12.Lianli Gao, Yaya Cheng, Qilong Zhang, Xing Xu, and Jingkuan Song. Feature space targeted attacks by statistic alignment. In IJCAI, 2021.
  13. 13.Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. In Yoshua Bengio and Yann LeCun (eds.), ICLR, 2015.
  14. 14.Lars Kai Hansen and Peter Salamon. Neural network ensembles. IEEE Trans. Pattern Anal. Mach. Intell., 1990.
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  16. 16.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In CVPR, 2018.
  17. 17.Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks. In CVPR, 2017.
  18. 18.Nathan Inkawhich, Wei Wen, Hai (Helen) Li, and Yiran Chen. Feature space perturbations yield more transferable adversarial examples. In CVPR, 2019.
  19. 19.Nathan Inkawhich, Kevin J. Liang, Binghui Wang, Matthew Inkawhich, Lawrence Carin, and Yiran Chen. Perturbing across the feature hierarchy to improve standard and strict blackbox attack transferability. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (eds.), NeurIPS, 2020.
  20. 20.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  21. 21.Jonathan Krause, Jia Deng, Michael Stark, and Li Fei-Fei. Collecting a large-scale dataset of finegrained cars. 2013.
  22. 22.Alex Krizhevsky. Learning multiple layers of features from tiny images. 2009.
  23. 23.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In NeurPIS, 2012.
  24. 24.Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. Adversarial examples in the physical world. In ICLR, 2017a.
  25. 25.Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio. Adversarial machine learning at scale. In ICLR, 2017b.
  26. 26.Xiaodan Li, Jinfeng Li, Yuefeng Chen, Shaokai Ye, Yuan He, Shuhui Wang, Hang Su, and Hui Xue. QAIR: practical query-efficient black-box attacks for image retrieval. In CVPR, 2021.
  27. 27.Yanpei Liu, Xinyun Chen, Chang Liu, and Dawn Song. Delving into transferable adversarial examples and black-box attacks. In ICLR, 2017.
  28. 28.Ye Liu, Yaya Cheng, Lianli Gao, Xianglong Liu, Qilong Zhang, and Jingkuan Song. Practical evaluation of adversarial robustness via adaptive auto attack. In CVPR, 2022.
  29. 29.Yantao Lu, Yunhan Jia, Jianyu Wang, Bai Li, Weiheng Chai, Lawrence Carin, and Senem Velipasalar. Enhancing cross-task black-box transferability of adversarial examples with dispersion reduction. In CVPR, 2020.
  30. 30.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In ICLR, 2018.
  31. 31.Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew B. Blaschko, and Andrea Vedaldi. Fine-grained visual classification of aircraft. volume abs/1306.5151, 2013.
  32. 32.Xiaofeng Mao, Yuefeng Chen, Shuhui Wang, Hang Su, Yuan He, and Hui Xue. Composite adversarial attacks. In AAAI, 2021.
  33. 33.Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard. Deepfool: A simple and accurate method to fool deep neural networks. In CVPR, 2016.
  34. 34.Konda Reddy Mopuri, Utkarsh Ojha, Utsav Garg, and R. Venkatesh Babu. NAG: network for adversary generation. In CVPR, 2018.
  35. 35.Muzammal Naseer, Salman H. Khan, Muhammad Haris Khan, Fahad Shahbaz Khan, and Fatih Porikli. Cross-domain transferability of adversarial perturbations. In NeurPIS, 2019.
  36. 36.Muzammal Naseer, Salman H. Khan, Munawar Hayat, Fahad Shahbaz Khan, and Fatih Porikli. A self-supervised approach for adversarial robustness. In CVPR, 2020.
  37. 37.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco Bo Wu, and Andrew Y. Ng. Reading digits in natural images with unsupervised feature learning. 2011.
  38. 38.Nicolas Papernot, Patrick D. McDaniel, and Ian J. Goodfellow. Transferability in machine learning: from phenomena to black-box attacks using adversarial samples. CoRR, abs/1605.07277, 2016.
  39. 39.Omid Poursaeed, Isay Katsman, Bicheng Gao, and Serge J. Belongie. Generative adversarial perturbations. In CVPR, 2018.
  40. 40.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li. Imagenet large scale visual recognition challenge. IJCV, 2015.
  41. 41.Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, and Michael K. Reiter. Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In SIGSAC, 2016.
  42. 42.Yucheng Shi, Siyu Wang, and Yahong Han. Curls & whey: Boosting black-box adversarial attacks. In CVPR, 2019.
  43. 43.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In Yoshua Bengio and Yann LeCun (eds.), ICLR, 2015.
  44. 44.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In Yoshua Bengio and Yann LeCun (eds.), ICLR, 2014.
  45. 45.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In CVPR, 2016.
  46. 46.Hugo Touvron, Alexandre Sablayrolles, Matthijs Douze, Matthieu Cord, and Herve J egou. Grafit: Learning fine-grained image representations with coarse labels. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 874–884, 2021.
  47. 47.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The Caltech-UCSD Birds-200-2011 Dataset. Technical report, California Institute of Technology, 2011.
  48. 48.Xiaosen Wang, Xuanran He, Jingdong Wang, and Kun He. Admix: Enhancing the transferability of adversarial attacks. In ICCV, 2021.
  49. 49.Xiaosong Wang, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M. Summers. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In CVPR, 2017.
  50. 50.Dongxian Wu, Yisen Wang, Shu-Tao Xia, James Bailey, and Xingjun Ma. Skip connections matter: On the transferability of adversarial examples generated with resnets. In ICLR, 2020a.
  51. 51.Weibin Wu, Yuxin Su, Xixian Chen, Shenglin Zhao, Irwin King, Michael R. Lyu, and Yu-Wing Tai. Boosting the transferability of adversarial samples via attention. In CVPR, 2020b.
  52. 52.Cihang Xie, Zhishuai Zhang, Yuyin Zhou, Song Bai, Jianyu Wang, Zhou Ren, and Alan L. Yuille. Improving transferability of adversarial examples with input diversity. In CVPR, 2019.
  53. 53.Kaidi Xu, Gaoyuan Zhang, Sijia Liu, Quanfu Fan, Mengshu Sun, Hongge Chen, Pin-Yu Chen, Yanzhi Wang, and Xue Lin. Adversarial t-shirt! evading person detectors in a physical world. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (eds.), ECCV, 2020.
  54. 54.Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. How transferable are features in deep neural networks? In Zoubin Ghahramani, Max Welling, Corinna Cortes, Neil D. Lawrence, and Kilian Q. Weinberger (eds.), NeurIPS, 2014.
  55. 55.Qilong Zhang, Chaoning Zhang, Chaoqun Li, Jingkuan Song, Lianli Gao, and Heng Tao Shen. Practical no-box adversarial attacks with training-free hybrid image transformation. CoRR, abs/2203.04607, 2022.
  56. 56.Zhengyu Zhao, Zhuoran Liu, and Martha A. Larson. On success and simplicity: A second look at transferable targeted attacks. CoRR, abs/2012.11207, 2020.
  57. 57.Wen Zhou, Xin Hou, Yongjun Chen, Mengyun Tang, Xiangqi Huang, Xiang Gan, and Yong Yang. Transferable adversarial perturbations. In Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss (eds.), ECCV, 2018.

Citation

MLA
Zhang, Q., et al. “Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains”. arXiv, 2022, http://arxiv.org/abs/2201.11528v4.
APA
Zhang, Q., Li, X., Chen, Y., Song, J., Gao, L., He, Y., & Xue, H. (2022). Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains. arXiv. http://arxiv.org/abs/2201.11528v4
Chicago
Zhang, Q., X. Li, Y. Chen, et al. 2022. “Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains”. arXiv. http://arxiv.org/abs/2201.11528v4.
Harvard
Zhang, Q. et al. (2022) “Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2201.11528v4.
Vancouver
1. Zhang Q, Li X, Chen Y, Song J, Gao L, He Y, Xue H (2022) Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains. arXiv

BibTeX

@article{zhang2022beyond,
  title = {Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains},
  author = {Zhang, Qilong and Li, Xiaodan and Chen, Yuefeng and Song, Jingkuan and Gao, Lianli and He, Yuan and Xue, Hui},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2201.11528v4},
  eprint = {2201.11528}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors