Adversarial Texture for Fooling Person Detectors in the Physical World
Zhanhao HuSiyuan HuangXiaopei ZhuFuchun SunBo ZhangXiaolin Hu
Proposes a generative method to produce repeatable adversarial textures for clothing that consistently fool physical-world person detectors across diverse viewing angles and fabric deformations.
Automated surveillance and security systems increasingly rely on deep neural networks to detect people in real time. However, these systems can be fooled by physical adversarial examples, such as specially designed printed patches placed on clothing. Existing patch-based attacks suffer from significant performance drops whenever camera angles shift or when cloth wrinkles distort the pattern, leaving major gaps in evaluating the true security vulnerabilities of computer vision systems.
The article demonstrates that an expandable adversarial pattern, termed Adversarial Texture (AdvTexture), can be synthesized in arbitrary dimensions to reliably evade person detectors from multiple viewing angles. The objective is to establish whether any local section of an adversarial pattern printed on regular garments can suppress human detections across varied real-world orientations.
To craft these textures, the authors developed a two-stage method known as Toroidal-Cropping-based Expandable Generative Attack (TC-EGA). The framework first trains a fully convolutional neural network to generate textures from spatial latent variables, ensuring translation invariance. In the second stage, it optimizes a seamless, repeatable local latent unit using a toroidal cropping technique to allow infinite tiling without losing effectiveness. The researchers evaluated the technique both digitally on standard benchmark datasets and physically by manufacturing real garments—including T-shirts, skirts, and dresses—tested on human subjects across indoor and outdoor settings.
The findings confirm that AdvTexture significantly outperforms existing attack methods across diverse conditions. In physical evaluations targeting the YOLOv2 detector, AdvTexture lowered detection average precision to 0.359, whereas prior patch methods achieved near 1.0 (indicating complete failure when the subject turned). At frontal and back angles (0° and 180°), the mean attack success rate approached 1.0, and overall evasion remained substantially higher across all angles compared to baseline patches. Garments with larger surface coverage, such as dresses, achieved the highest mean attack success rates (0.893), compared to T-shirts (0.771) and skirts (0.287). The attacks also generalized well across other detectors, including Faster R-CNN (0.930 success rate) and Mask R-CNN (0.855 success rate).
These results demonstrate a serious physical vulnerability in widely deployed security and surveillance infrastructures. Patch attacks can no longer be assumed to fail simply because a target is moving, turning, or wearing non-rigid clothing. Security teams and system designers must recognize that adversaries can disguise themselves from standard automated vision pipelines using everyday wearable textiles.
To counter these vulnerabilities, organizations deploying automated vision security should implement robust defense mechanisms, such as input denoising, feature squeezing, and model ensemble strategies. Future development should also explore multi-model adversarial training to anticipate and neutralize transferable texture attacks.
The primary limitation noted in the article is that textures optimized against a single target detector experience reduced transferability against different architectures without further multi-model training. In addition, detection evasion decreases at steep side profiles (around 90° and 270°) and at extended distances where camera coverage is limited. Overall confidence in the findings is high for near-to-moderate distances across standard viewing angles.
- Paper: Adversarial Patch, Tom B. Brown et al. (2017). Introduces printable, physical adversarial patches designed to fool image vision systems, which the source builds upon to develop wearable adversarial clothing textures.
- Paper: Synthesizing Robust Adversarial Examples, Anish Athalye et al. (2017). Pioneers the Expectation Over Transformation framework to synthesize robust physical-world adversarial examples invariant to 3D rotations, angles, and lighting.
- Paper: Adversarial examples in the physical world, Alexey Kurakin et al. (2016). Establishes the foundational empirical methodology demonstrating that adversarial examples survive printing and physical camera recapture.
- Paper: ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness, Robert Geirhos et al. (2018). Analyzes the fundamental texture bias in convolutional neural networks, providing the theoretical motivation for crafting adversarial textures to fool visual detectors.
- Paper: Improving Transferability of Adversarial Examples With Input Diversity, Cihang Xie et al. (2018). Demonstrates applying input transformations during gradient optimization to prevent overfitting and improve attack robustness across diverse viewing environments.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). Explains the foundational mechanics of generating adversarial perturbations via gradient-based optimization in deep neural networks.
- Paper: Towards Deep Learning Models Resistant to Adversarial Attacks, Aleksander Madry et al. (2017). Formulates adversarial attack generation through projected gradient descent under a robust optimization saddle-point framework.
- Paper: Shadows can be Dangerous: Stealthy and Effective Physical-world Adversarial Attack by Natural Phenomenon, Yiqi Zhong et al. (2022). Extends physical-world adversarial evasion beyond wearable textures to non-invasive, naturally occurring geometric shadows for camera deception.
- Paper: Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation, Zeyu Qin et al. (2022). Broadens the transferability of adversarial perturbations across unknown black-box vision models by optimizing over flat loss landscapes.
- Paper: Visual Adversarial Examples Jailbreak Aligned Large Language Models, Xiangyu Qi et al. (2024). Explores the frontier of visual adversarial attacks by applying image perturbations to break the safety alignment of multimodal vision-language models.
