Black-box Adversarial Attacks with Limited Queries and Information
Andrew IlyasLogan EngstromAnish AthalyeJessy Lin
Develops practical black-box adversarial attack methods that fool image classifiers under strict query and feedback constraints, proving their real-world threat by successfully deceiving the Google Cloud Vision API.
Commercial and proprietary artificial intelligence systems frequently rely on "black-box" security assumptions, where internal model parameters, training data, and gradients remain hidden from users. However, standard black-box attack evaluations assume an adversary can make infinite queries and receive complete probability distributions over all possible classes. In commercial applications, systems impose query limits, introduce monetary costs per query, restrict output to partial scores, or provide only sorted class labels without confidence metrics. Understanding whether neural networks remain vulnerable under these strict operational constraints is critical for assessing the actual security of deployed machine learning systems.
The article demonstrates that targeted adversarial attacks remain highly effective even under realistic real-world restrictions. It introduces new algorithms designed to fool classifiers across three constrained threat models: query-limited, partial-information, and label-only settings.
To evaluate vulnerability without internal model access, the authors adapted Natural Evolution Strategies (NES) to estimate gradients using derivative-free optimization over random Gaussian noise. For partial-information settings, the authors developed an optimization algorithm that begins with an image of the target class and alternates between projected gradient descent to maximize target confidence and backtracking line searches to blend in the source image. For label-only settings, where classifiers provide no probability scores, the method uses random local noise perturbations to calculate a proxy score based on how consistently the label appears. These techniques were systematically tested on the standard InceptionV3 image classification benchmark using 1,000 test images (100 for label-only tests) and validated on a live commercial platform, the Google Cloud Vision API.
The experimental findings show that restrictive interfaces fail to prevent targeted adversarial manipulation. In the query-limited setting, the attack achieved a 99.2% targeted success rate with a median of 11,550 queries—reducing the query cost by two to three orders of magnitude compared to prior gradient estimation approaches. In the partial-information setting where only top-1 probabilities are exposed, the attack achieved a 93.6% success rate requiring a median of 49,624 queries. In the label-only setting where no scores are returned, the attack reached a 90.0% success rate with a median of 2.7 million queries. Finally, the authors successfully executed a targeted attack against the Google Cloud Vision API, converting an image of skiers into an adversarial image that appeared visually unchanged to humans but was classified by the API as a dog.
These results demonstrate that commercial machine learning models are vulnerable to targeted manipulation despite strict output filtering, paywalls, or rate-limiting. Limiting output feedback does increase the query budget required for an attack, but it does not provide security against determined adversaries. Security and risk teams should not rely on black-box opacity or simple output stripping as a sufficient defense, and future defensive strategies must focus on model robustness rather than superficial interface restrictions. While the empirical results firmly establish this vulnerability across both benchmark models and commercial APIs, decision-makers should note that the label-only setting requires substantial query volumes that may be detectable under strict traffic monitoring or rate limiting.
- Paper: ZOO: Zeroth Order Optimization Based Black-box Attacks to Deep Neural Networks without Training Substitute Models, Pin-Yu Chen et al. (2017). This paper establishes gradient estimation via zeroth-order optimization as a standard technique for black-box score-based attacks, which the source paper directly improves upon for query efficiency and partial information.
- Paper: Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning Models, Wieland Brendel et al. (2017). This work introduces the decision-based boundary attack setting that relies strictly on top-1 label outputs, defining the label-only threat model that the source paper optimizes.
- Paper: Practical Black-Box Attacks against Machine Learning, Nicolas Papernot et al. (2017). This foundational study demonstrates practical black-box attacks against commercial cloud APIs using substitute models, framing the real-world black-box threat settings addressed in the source.
- Paper: Delving into Transferable Adversarial Examples and Black-box Attacks, Yanpei Liu et al. (2016). This research analyzes black-box attack feasibility and ensemble transferability on large-scale datasets, providing essential context for black-box evaluations on ImageNet.
- Paper: Towards Evaluating the Robustness of Neural Networks, Nicholas Carlini et al. (2016). This seminal paper formulates powerful optimization objectives for crafting adversarial perturbations, establishing the baseline loss formulations adapted for black-box numerical optimization.
- Paper: Synthesizing Robust Adversarial Examples, Anish Athalye et al. (2017). This work introduces Expectation Over Transformation and natural evolution strategies for robust optimization under uncertainty, serving as core algorithmic tools for query-efficient gradient estimation.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, Aleksander Madry et al. (2017). This paper formalizes projected gradient descent and first-order adversarial optimization against deep networks, defining the optimization standards and perturbation constraints assumed throughout the source.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This foundational text explains the linear nature of adversarial vulnerabilities in neural networks and introduces fundamental gradient sign methods for generating adversarial perturbations.
- Paper: Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks, Francesco Croce et al. (2020). This paper integrates score-based black-box attacks and decision-based concepts into a standardized, parameter-free evaluation suite (AutoAttack) to benchmark adversarial robustness reliably.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). This work extends the foundational understanding of adversarial examples by demonstrating that vulnerabilities exploited by black-box attacks stem from predictive, non-robust features learned by standard classifiers.
- Paper: Improving Transferability of Adversarial Examples With Input Diversity, Cihang Xie et al. (2018). This study advances black-box transferability by incorporating input diversity transformations during generation, building on the practical limitations of black-box adversaries.
- Paper: Certified Adversarial Robustness via Randomized Smoothing, Jeremy M Cohen et al. (2019). This paper introduces randomized smoothing to provide provable, certified robustness guarantees against continuous adversarial perturbations, addressing the vulnerabilities exposed by black-box and gradient-free attacks.
- Paper: Theoretically Principled Trade-off between Robustness and Accuracy, Hongyang Zhang et al. (2019). This work establishes a theoretically grounded defense framework (TRADES) that balances natural classification performance against robustness to adversarial optimization.
- Paper: Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment, Di Jin et al. (2019). This research applies black-box, score-based adversarial attack principles to the discrete domain of natural language processing and transformer models like BERT.
