Countering Adversarial Images using Input Transformations
Chuan GuoMayank RanaMoustapha CisseLaurens van der Maaten
Demonstrates that applying non-differentiable input transformations like image quilting and total variance minimization can effectively defend convolutional networks against strong adversarial attacks on ImageNet without requiring complex model modifications.
As deep learning systems are increasingly deployed in security-critical environments such as autonomous driving and medical imaging, their susceptibility to adversarial attacks poses a serious operational risk. Adversaries can introduce small, imperceptible alterations to input images that cause state-of-the-art neural networks to make incorrect classifications. Existing defense strategies generally fall into two categories: model-specific approaches that modify neural networks but often break when attackers adapt their strategies, and model-agnostic input defenses that historically proved too simple to be fully effective.
The article evaluates whether preprocessing input images with specialized transformations can effectively remove adversarial perturbations before images reach the classifier. The objective is to establish practical, model-agnostic defenses that maintain high classification accuracy on normal images while neutralizing diverse attack strategies, even when attackers have full knowledge of the classifier's internal architecture and parameters.
To assess this, the authors tested five image transformation techniques: cropping and rescaling, bit-depth reduction, JPEG compression, total variation minimization (a method that reconstructs smooth images after randomly dropping pixels), and image quilting (which synthesizes images using small, clean reference patches). These defenses were evaluated on the standard ImageNet dataset using deep residual networks against four prominent attack methods across both black-box scenarios (where the adversary has no access to the model) and gray-box scenarios (where the adversary knows the network parameters but not the specific preprocessing defense).
The evaluation revealed several key findings. First, training the neural networks directly on transformed images substantially improved defense performance compared to merely applying transformations at test time. Second, image quilting and total variation minimization served as the most resilient defenses; under strong black-box attacks, image quilting neutralized 80% to 90% of threats. Third, combining multiple defenses—specifically ensembling cropping, total variation minimization, and quilting across modern architectures—achieved an accuracy of approximately 71%, suffering at most a 6% performance drop under heavy attack. Fourth, in gray-box evaluations against iterative attacks such as DeepFool, transformation-based defenses achieved over 50% accuracy, outperforming prior state-of-the-art ensemble adversarial training methods by 18 to 24 times.
These findings indicate that effective defenses must rely on transformations that are non-differentiable and inherently randomized. Randomness prevents adversaries from mathematically calculating exact perturbations, forcing them to solve a significantly harder problem of fooling a broad distribution of potential image reconstructions. This offers an accessible, scalable way to harden existing computer vision pipelines without requiring complex, attack-specific retraining protocols.
Organizations deploying vision systems should consider adopting randomized preprocessing pipelines—particularly combining spatial cropping, total variation minimization, and quilting—and retraining core classifiers on these transformed inputs. While effective against current attack vectors, these defenses assume the adversary does not know the specific real-time preprocessing configuration. Future engineering work should evaluate how these defenses withstand adaptive attacks designed specifically to bypass randomized transformations, and investigate extending similar transformation defenses to other domains such as speech and audio processing.
- Paper: Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks, Weilin Xu et al. (2017). Introduces feature squeezing through input-level transformations like bit-depth reduction and spatial filtering to counter adversarial examples, laying the groundwork for the preprocessing defense strategies examined in the source.
- Paper: Towards Evaluating the Robustness of Neural Networks, Nicholas Carlini et al. (2016). Establishes standard optimization-based adversarial attacks across distance metrics, which serve as foundational threat models evaluated against the input transformation defenses.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). Explains the foundational linearity hypothesis and fast gradient sign method for generating adversarial examples, which forms the basis of the attacks evaluated in the source.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, Aleksander Madry et al. (2017). Frames adversarial robustness as a robust minimax optimization problem and formalizes strong multi-step projected gradient descent attacks tested against the defenses.
- Paper: Adversarial Machine Learning at Scale, Alexey Kurakin et al. (2016). Demonstrates the behavior and scaling of adversarial attacks and defenses on large-scale ImageNet classifiers, setting up the primary benchmark setting used in the source.
- Paper: Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods, Nicholas Carlini et al. (2017). Analyzes why ad-hoc detection and preprocessing defenses often fail against adaptive adversaries, providing essential context for designing robust input transformation defenses.
- Paper: DeepFool: A Simple and Accurate Method to Fool Deep Neural Networks, Seyed-Mohsen Moosavi-Dezfooli et al. (2015). Develops the DeepFool algorithm to efficiently compute minimal adversarial perturbations, providing a core attack benchmark used to evaluate input transformations.
- Paper: Intriguing properties of neural networks, Christian Szegedy et al. (2014). First discovered the existence of adversarial examples and the discontinuous input-output mappings of deep networks that necessitate transformation-based defenses.
- Paper: Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples, Anish Athalye et al. (2018). Directly critiques non-differentiable and randomized input transformation defenses like quilting and total variance minimization by proposing Backward Pass Differentiable Approximation (BPDA) and Expectation Over Transformation to bypass them.
- Paper: Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks, Francesco Croce et al. (2020). Develops the parameter-free AutoAttack benchmark to rigorously evaluate and circumvent defenses that rely on gradient masking or input transformations.
- Paper: Certified Adversarial Robustness via Randomized Smoothing, Jeremy M Cohen et al. (2019). Transitions from empirical, heuristic input transformation defenses to provable, certified adversarial robustness via randomized smoothing.
- Paper: Adversarial Examples Are Not Bugs, They Are Features, Andrew Ilyas et al. (2019). Provides a fundamental feature-level explanation for why heuristic input transformations struggle by showing that adversarial examples stem from non-robust features inherent to standard training datasets.
- Paper: Robustness May Be at Odds with Accuracy, Dimitris Tsipras et al. (2018). Investigates the fundamental trade-off between standard accuracy and adversarial robustness observed when training classifiers on transformed and adversarial inputs.
- Paper: Benchmarking Neural Network Robustness to Common Corruptions and Perturbations, Dan Hendrycks et al. (2019). Broadens the evaluation of image transformations and corruptions beyond targeted adversarial noise to standardized common corruptions and perturbations.
- Paper: AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty, Dan Hendrycks et al. (2019). Builds upon input transformation concepts by employing diverse augmentation chains and consistency regularization to enhance out-of-distribution robustness.
