Beyond ImageNet Attack: Towards Crafting Adversarial Examples for Black-box Domains
Qilong ZhangXiaodan LiYuefeng ChenJingkuan SongLianli GaoYuan HeHui Xue
Develops a generative framework that disrupts low-level image features using only ImageNet knowledge, enabling transferable black-box attacks against models deployed on completely unknown target domains.
Deep neural networks are widely used in mission-critical vision applications, yet they remain vulnerable to adversarial attacks—subtle, engineered perturbations to input images that cause models to make incorrect predictions. Previous security assessments primarily assumed that attackers have access to the target model's training data or can repeatedly probe the deployed system with queries. However, real-world systems are rarely accessible for frequent querying, and their underlying training datasets are usually kept private. This creates an urgent operational need to evaluate whether vision systems can be compromised under strict black-box conditions, where attackers have zero knowledge of the target data, model architecture, or task.
The main objective of the article is to demonstrate that attackers can craft highly transferable adversarial examples for unknown black-box domains using only a substitute model trained on standard, publicly available ImageNet data. It introduces the Beyond ImageNet Attack framework and evaluates its ability to fool diverse, unknown target classifiers across eight distinct image classification tasks.
To conduct this evaluation, the researchers trained a generative model on ImageNet to learn adversarial perturbations that disrupt general, low-level visual features in intermediate network layers rather than task-specific final layers. The study tested the approach across four substitute architectures and evaluated attack success against seven target domains, comprising four coarse-grained datasets and three fine-grained datasets. To bridge the gap between source and target domains, the authors introduced two enhancements: a random normalization module to simulate varying data distributions and a domain-agnostic attention module to focus perturbations on core object features.
The analysis produced several key findings. First, the proposed method substantially outperforms existing state-of-the-art transfer attacks across all tested domains. In fine-grained classification tasks, the domain-agnostic attention variant reduced target model accuracy to an average of 36.42%, outperforming the strongest baseline by roughly 25.9 percentage points. In coarse-grained tasks, the random normalization variant reduced target accuracy to an average of 57.30%, beating baseline methods by 7.71 percentage points. Second, attacking intermediate and shallow network layers proved significantly more effective for cross-domain transfer than attacking deep, task-specific layers. Third, combining substitute models into an ensemble further increased attack strength, while standard data augmentation failed to provide similar transfer benefits.
These findings carry critical security and risk implications. They demonstrate that withholding training data, hiding model architectures, and restricting query access provide a false sense of security. Commercially deployed vision systems can be fooled in real time with a single forward pass by adversaries leveraging public datasets. Consequently, security assessments that assume isolated models are safe from transfer attacks significantly underestimate operational vulnerabilities in sensitive environments such as autonomous navigation, healthcare diagnostics, and automated surveillance.
Based on these results, machine learning practitioners and security teams should immediately incorporate cross-domain adversarial testing into their model validation pipelines rather than relying solely on standard robustness evaluations. Defenses must focus on hardening low-level feature representations and implementing robust input pre-processing. Combining random normalization and domain-agnostic attention should be tailored cautiously, as the article observed that using both simultaneously can produce mixed results depending on the underlying network backbone.
The primary limitation of the study is its focus on image classification benchmarks using standard convolutional backbones, leaving potential impacts on other vision tasks—such as object detection and segmentation—untested. Furthermore, generating effective cross-domain attacks requires training on a sufficiently large and diverse source dataset like ImageNet; models trained on smaller datasets transfer poorly. Nevertheless, given the consistent outperformance across multiple benchmarks and repeated runs with low standard deviations, confidence in the demonstrated vulnerabilities remains very high.
- Paper: Delving into Transferable Adversarial Examples and Black-box Attacks, Yanpei Liu et al. (2016). This foundational paper establishes how adversarial examples transfer across diverse deep learning architectures on ImageNet and introduces multi-model ensembles to enhance black-box transferability.
- Paper: Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples, Nicolas Papernot et al. (2016). This work introduces the core substitute-model paradigm for black-box transfer attacks that the source paper seeks to generalize to unknown domains without query access.
- Paper: Improving Transferability of Adversarial Examples With Input Diversity, Cihang Xie et al. (2018). This paper demonstrates how applying input diversity and transformations during optimization prevents overfitting to substitute networks and improves transferability.
- Paper: Universal Adversarial Perturbations, Seyed-Mohsen Moosavi-Dezfooli et al. (2016). This study introduces data-agnostic universal adversarial perturbations that generalize across images and network architectures, laying groundwork for cross-domain attack formulations.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This seminal work provides the foundational mathematical explanation for why adversarial perturbations transfer across models and presents the Fast Gradient Sign Method.
- Paper: Intriguing properties of neural networks, Christian Szegedy et al. (2014). This foundational study discovers the existence of adversarial examples and demonstrates their general cross-model transferability in deep neural networks.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). This paper provides essential theoretical insight into how deep neural networks rely on non-robust visual features that transfer across models and distributions.
- Paper: Learning Robust Global Representations by Penalizing Local Predictive Power, Haohan Wang et al. (2019). This work analyzes how early convolutional layers capture localized cues and domain-shift vulnerabilities, motivating the source paper's focus on attacking intermediate representations.
- Paper: Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition, Zexin Li et al. (2023). This paper extends cross-model black-box transfer attacks to specialized face recognition domains by leveraging multi-task auxiliary representations.
