Adversarial Texture for Fooling Person Detectors in the Physical World

Zhanhao HuSiyuan HuangXiaopei ZhuFuchun SunBo ZhangXiaolin Hu

article2022CVPR176 citations

Proposes a generative method to produce repeatable adversarial textures for clothing that consistently fool physical-world person detectors across diverse viewing angles and fabric deformations.

Listen

Automated surveillance and security systems increasingly rely on deep neural networks to detect people in real time. However, these systems can be fooled by physical adversarial examples, such as specially designed printed patches placed on clothing. Existing patch-based attacks suffer from significant performance drops whenever camera angles shift or when cloth wrinkles distort the pattern, leaving major gaps in evaluating the true security vulnerabilities of computer vision systems.

The article demonstrates that an expandable adversarial pattern, termed Adversarial Texture (AdvTexture), can be synthesized in arbitrary dimensions to reliably evade person detectors from multiple viewing angles. The objective is to establish whether any local section of an adversarial pattern printed on regular garments can suppress human detections across varied real-world orientations.

To craft these textures, the authors developed a two-stage method known as Toroidal-Cropping-based Expandable Generative Attack (TC-EGA). The framework first trains a fully convolutional neural network to generate textures from spatial latent variables, ensuring translation invariance. In the second stage, it optimizes a seamless, repeatable local latent unit using a toroidal cropping technique to allow infinite tiling without losing effectiveness. The researchers evaluated the technique both digitally on standard benchmark datasets and physically by manufacturing real garments—including T-shirts, skirts, and dresses—tested on human subjects across indoor and outdoor settings.

The findings confirm that AdvTexture significantly outperforms existing attack methods across diverse conditions. In physical evaluations targeting the YOLOv2 detector, AdvTexture lowered detection average precision to 0.359, whereas prior patch methods achieved near 1.0 (indicating complete failure when the subject turned). At frontal and back angles (0° and 180°), the mean attack success rate approached 1.0, and overall evasion remained substantially higher across all angles compared to baseline patches. Garments with larger surface coverage, such as dresses, achieved the highest mean attack success rates (0.893), compared to T-shirts (0.771) and skirts (0.287). The attacks also generalized well across other detectors, including Faster R-CNN (0.930 success rate) and Mask R-CNN (0.855 success rate).

These results demonstrate a serious physical vulnerability in widely deployed security and surveillance infrastructures. Patch attacks can no longer be assumed to fail simply because a target is moving, turning, or wearing non-rigid clothing. Security teams and system designers must recognize that adversaries can disguise themselves from standard automated vision pipelines using everyday wearable textiles.

To counter these vulnerabilities, organizations deploying automated vision security should implement robust defense mechanisms, such as input denoising, feature squeezing, and model ensemble strategies. Future development should also explore multi-model adversarial training to anticipate and neutralize transferable texture attacks.

The primary limitation noted in the article is that textures optimized against a single target detector experience reduced transferability against different architectures without further multi-model training. In addition, detection evasion decreases at steep side profiles (around 90° and 270°) and at extended distances where camera coverage is limited. Overall confidence in the findings is high for near-to-moderate distances across standard viewing angles.

arXiv: 2203.03373
Cover for Adversarial Texture for Fooling Person Detectors in the Physical World

Abstract

Nowadays, cameras equipped with AI systems can capture and analyze images to detect people automatically. However, the AI system can make mistakes when receiving deliberately designed patterns in the real world, i.e., physical adversarial examples. Prior works have shown that it is possible to print adversarial patches on clothes to evade DNN-based person detectors. However, these adversarial examples could have catastrophic drops in the attack success rate when the viewing angle (i.e., the camera's angle towards the object) changes. To perform a multi-angle attack, we propose Adversarial Texture (AdvTexture). AdvTexture can cover clothes with arbitrary shapes so that people wearing such clothes can hide from person detectors from different viewing angles. We propose a generative method, named Toroidal-Cropping-based Expandable Generative Attack (TC-EGA), to craft AdvTexture with repetitive structures. We printed several pieces of cloth with AdvTexture and then made T-shirts, skirts, and dresses in the physical world. Experiments showed that these clothes could fool person detectors in the physical world.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methods
  • 3.1. Adversarial Patch Generator
  • 3.1.1. The Adversary Objective Function
  • 3.1.2. The Information Objective function
  • 3.2. Toroidal-Cropping-based Expandable Generative Attack
  • 3.2.1. Stage One: Train an Expandable Generator
  • 3.2.2. Stage Two: Find the Best Latent Pattern
  • 4. Experiment settings
  • 4.1. Subjects
  • 4.2. Dataset
  • 4.3. Baseline Methods
  • 4.4. Implementation Details
  • 5. Results
  • 5.1. Patch-Based Attack in the Digital World
  • 5.2. Attack in the Physical World
  • 6. Conclusions
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Expandable adversarial texture for multi-angle person evasion

    model/method

    The paper introduces Adversarial Texture (AdvTexture): a texture that can be generated at arbitrary spatial sizes and printed over clothing of arbitrary shape. Unlike a single adversarial patch, AdvTexture is designed so that approximately every local region extracted from the cloth is adversarially effective. Consequently, a person wearing the clothing can remain difficult to detect even when a camera captures different, deformed, or partial regions of the clothing at different viewing angles.

    This design addresses the segment-missing problem. A single patch becomes ineffective when the camera captures only part of it, and tightly tiling conventional patches can expose boundaries or combine ineffective portions of different patch units. AdvTexture instead uses a spatially repetitive structure whose local regions retain adversarial effectiveness.

  2. Knowl 2 — Generative distribution and mutual-information training objective

    equation

    Let τ~\tilde{\tau} denote an adversarial patch, U(τ~)U(\tilde{\tau}) its scalar energy, and padvp_{\mathrm{adv}} the desired distribution of adversarial patches. The paper defines

    padv(τ~)=exp⁡[−U(τ~)]ZU,ZU=∫exp⁡[−U(τ~)]dτ~.p_{\mathrm{adv}}(\tilde{\tau})=\frac{\exp[-U(\tilde{\tau})]}{Z_U},\qquad Z_U=\int \exp[-U(\tilde{\tau})]d\tilde{\tau}.

    Because direct sampling from this distribution requires the generally intractable partition function ZUZ_U, a generator GφG_\varphi with parameters φ\varphi maps a latent variable z∼N(0,I)z\sim\mathcal{N}(0,I) to a patch τ~=Gφ(z)\tilde{\tau}=G_\varphi(z). The resulting generator distribution is

    qφ(τ~)=∫δ ⁣(τ~−Gφ(z))pz(z) dz,q_\varphi(\tilde{\tau})=\int \delta\!\left(\tilde{\tau}-G_\varphi(z)\right)p_z(z)\,dz,

    where pzp_z is the standard-normal density and δ\delta is the Dirac delta function. Minimizing KL(qφ∥padv)\mathrm{KL}(q_\varphi\|p_{\mathrm{adv}}) is implemented using an auxiliary scalar network Tω(τ~,z)T_\omega(\tilde{\tau},z) and the objective

    min⁡φ,ω  Eτ~∼qφ[U(τ~)]−Iφ,ωJSD(τ~,z),\min_{\varphi,\omega}\;\mathbb{E}_{\tilde{\tau}\sim q_\varphi}[U(\tilde{\tau})]-I^{\mathrm{JSD}}_{\varphi,\omega}(\tilde{\tau},z),

    where

    Iφ,ωJSD=E(τ~,z)∼qφτ~,z[−sp⁡ ⁣(−Tω(τ~,z))]−Eτ~∼qφ,  z′∼pz[sp⁡ ⁣(Tω(τ~,z′))],I^{\mathrm{JSD}}_{\varphi,\omega}=\mathbb{E}_{(\tilde{\tau},z)\sim q^{\tilde{\tau},z}_\varphi}\left[-\operatorname{sp}\!\left(-T_\omega(\tilde{\tau},z)\right)\right]-\mathbb{E}_{\tilde{\tau}\sim q_\varphi,\;z'\sim p_z}\left[\operatorname{sp}\!\left(T_\omega(\tilde{\tau},z')\right)\right],

    and sp⁡(t)=log⁡(1+et)\operatorname{sp}(t)=\log(1+e^t) is the softplus function. The first term lowers detector-related energy, while the negative mutual-information term encourages different latent variables to generate different patches rather than collapsing to one output.

  3. Knowl 3 — Detector-aware patch energy with physical transformations and smoothness

    model/method

    For an original training image xx, a generated patch τ~\tilde{\tau}, and a detector transformation-and-attachment process MM, the adversarial component of the energy is

    Uobj=Ex,M[f ⁣(M(x,τ~))],U_{\mathrm{obj}}=\mathbb{E}_{x,M}\left[f\!\left(M(x,\tilde{\tau})\right)\right],

    where ff denotes the confidence scores of the person bounding boxes predicted by the target detector. The process MM randomly attaches a generated patch to a person according to predicted boxes and applies physical transformations including changes in scale, contrast, brightness, additive noise, and Thin Plate Spline deformation. Minimizing UobjU_{\mathrm{obj}} reduces the confidence of predicted person boxes, making them more likely to be removed by confidence-based filtering.

    To favor printable, smoother textures, the method adds a total-variation penalty over neighboring pixels:

    UTV=∑i,j(∣τ~i,j−τ~i+1,j∣+∣τ~i,j−τ~i,j+1∣),U_{\mathrm{TV}}=\sum_{i,j}\left(\left|\tilde{\tau}_{i,j}-\tilde{\tau}_{i+1,j}\right|+\left|\tilde{\tau}_{i,j}-\tilde{\tau}_{i,j+1}\right|\right),

    where τ~i,j\tilde{\tau}_{i,j} is the pixel value at spatial location (i,j)(i,j) and the sum covers valid horizontal and vertical neighboring pairs. The complete energy is

    U(τ~)=1β(Uobj+αUTV),U(\tilde{\tau})=\frac{1}{\beta}\left(U_{\mathrm{obj}}+\alpha U_{\mathrm{TV}}\right),

    where α\alpha weights smoothness and β\beta rescales the combined energy.

  4. Knowl 4 — Fully convolutional expandable generator

    model/method

    The first stage of Toroidal-Cropping-based Expandable Generative Attack (TC-EGA) trains an all-convolutional generator. Every layer, including the layer receiving the latent variable, uses convolution with zero padding. The latent variable is a tensor of shape B×C×H×WB\times C\times H\times W, where BB is batch size, CC is the number of latent channels, and H,WH,W are spatial dimensions.

    Because convolution is translation invariant, a patch of spatial shape (w,h)(w,h) extracted around any location of the generated texture can be viewed as the output of a location-independent sub-generator Gw,hG_{w,h}. If the corresponding latent components have the same distribution at every location, then the extracted patch distribution is also location independent. Thus, training Gw,hG_{w,h} to approximate the adversarial patch distribution makes patches extracted from arbitrary positions approximately adversarially effective.

    The generator is trained using minimum latent dimensions Hmin⁡H_{\min} and Wmin⁡W_{\min}. Once trained, it can generate textures with any latent dimensions satisfying H≥Hmin⁡H\geq H_{\min} and W≥Wmin⁡W\geq W_{\min} simply by expanding the latent tensor spatially, without retraining the generator. This is the theoretical basis for producing large clothing-covering textures from a relatively small training configuration.

  5. Knowl 5 — Toroidal Cropping for optimizing a repeated latent pattern

    algorithm

    TC-EGA's second stage searches for a particularly effective latent pattern while keeping the trained fully convolutional generator fixed. Directly optimizing a large latent tensor is difficult and can produce discontinuities at crop boundaries, so the method optimizes only a small pattern and repeats it with wraparound continuity.

    Let zlocalz_{\mathrm{local}} have shape B×C×L×LB\times C\times L\times L. Its opposite horizontal edges and opposite vertical edges are identified, conceptually folding the pattern into a two-dimensional torus. Any required latent tensor can then be obtained by repeatedly tiling zlocalz_{\mathrm{local}} in both spatial directions and cropping the resulting periodic tensor. A crop that crosses a tile boundary is therefore continuous, equivalent to cropping across the identified edges of the torus.

    The optimization procedure is:

    1. Initialize zlocalz_{\mathrm{local}} with shape B×C×L×LB\times C\times L\times L.
    2. At each optimization step, use toroidal tiling and cropping to construct a sample zsamplez_{\mathrm{sample}} with the minimum generator input shape B×C×Hmin⁡×Wmin⁡B\times C\times H_{\min}\times W_{\min}.
    3. Generate a patch Gφ(zsample)G_\varphi(z_{\mathrm{sample}}) with the frozen generator.
    4. Apply the adversarial objective, including the detector simulation and physical transformations, and update zlocalz_{\mathrm{local}} with Adam.
    5. After optimization converges under the training schedule, tile the optimized local pattern to any desired latent spatial size and pass it through the generator to obtain the final AdvTexture.

    The paper also evaluates a pixel-space toroidal variant, but the proposed TC-EGA performs toroidal cropping in latent space after training the expandable generator.

  6. Knowl 6 — Training and physical implementation configuration

    experimental setup

    The experiments use the Inria Person dataset, with 614 training images and 288 test images. Target detectors are YOLOv2, YOLOv3, Faster R-CNN, and Mask R-CNN pretrained on MS COCO and restricted to the person class. Training-image detections are obtained with non-maximum suppression threshold 0.40.4; boxes with confidence below 0.50.5 for YOLOv2/YOLOv3 or below 0.750.75 for Faster R-CNN/Mask R-CNN are discarded, as are Faster R-CNN and Mask R-CNN boxes covering less than 0.16%0.16\% of the image.

    Adam is used in both TC-EGA stages. The first-stage generator is a seven-layer fully convolutional network with learning rate 0.0010.001, latent input shape B×128×9×9B\times128\times9\times9, and RGB output shape B×3×324×324B\times3\times324\times324. The second stage optimizes a local latent tensor of shape 1×128×4×41\times128\times4\times4 with learning rate 0.030.03; toroidal cropping expands it into samples of shape B×128×9×9B\times128\times9\times9.

    For physical evaluation, the generated textures are digitally printed on polyester cloth and tailored into T-shirts, skirts, and dresses. Baselines include the original AdvPatch and AdvTshirt attacks, tiled versions of each, repetitive random colors, and three ablations: EGA, which uses the expandable generator without latent optimization; TCA, which directly optimizes a 300×300300\times300 local texture with 150×150150\times150 toroidal crops; and RCA, which directly optimizes fixed large textures of 300×300300\times300 or 900×900900\times900 pixels using random 150×150150\times150 crops.

  7. Knowl 7 — Digital attack performance and component ablation

    data/table

    Digital evaluation attaches patches to images from the Inria test set and measures average precision (AP) of the detector's proposed boxes. The detector's predictions on clean images are used as ground truth, so lower AP indicates a stronger attack. Expandable means the method can produce arbitrary-sized textures; resampled means evaluation uses randomly extracted patches.

    Method AP Expandable Resampled
    Clean 1.000 No No
    Random 0.963 Yes Yes
    AdvPatch 0.352 No No
    AdvPatchTile 0.827 Yes Yes
    AdvTshirt 0.744^* No No
    AdvTshirtTile 0.844 Yes Yes
    TC-EGA 0.362 Yes Yes
    EGA 0.470 Yes Yes
    TCA 0.664 Yes Yes
    RCA_2× 0.606 No Yes
    RCA_6× 0.855 No Yes

    TC-EGA obtains the lowest AP among the methods evaluated with randomly resampled patches. EGA is weaker than TC-EGA, showing the benefit of optimizing the latent pattern after generator training. TCA is also weaker, indicating that latent-space toroidal optimization is more effective than directly optimizing pixels. The 900×900900\times900 RCA texture performs substantially worse than the 300×300300\times300 version, indicating difficulty in optimizing a large fixed texture. Tiling conventional patches greatly increases AP relative to their original single-patch forms, demonstrating the segment-missing problem. The asterisk on AdvTshirt denotes that it was trained on a different dataset.

  8. Knowl 8 — Physical-world evaluation protocol

    experimental setup

    Three subjects participated in the physical test set: their mean age was 24.024.0 years, the age range was 2121--2626, and the group contained two males and one female. Each subject wore an adversarial garment and slowly turned through a circle in front of a fixed camera positioned 1.381.38 m above the ground; the default person-to-camera distance was 22 m. For each subject and garment, one indoor video and one outdoor video were recorded, giving six videos per garment. Thirty-two frames were extracted from each video, yielding 192 manually labeled frames per garment.

    The physical evaluation uses recall-versus-confidence curves and attack success rate. For a confidence threshold, a detection is counted as correct if at least one predicted box above the threshold has intersection over union greater than 0.50.5 with the ground-truth person box. The attack success rate (ASR) is the fraction of images without a correct detection, and mean ASR (mASR) averages ASR over thresholds 0.1,0.2,…,0.90.1,0.2,\ldots,0.9.

  9. Knowl 9 — Physical attacks remain effective across viewing angles and detectors

    data/table

    On the physical YOLOv2 test set, the reported AP values are Random 1.0001.000, AdvPatch 0.9950.995, AdvPatchTile 0.9960.996, AdvTshirt 1.0001.000, AdvTshirtTile 0.9520.952, and TC-EGA 0.3590.359. Lower AP and lower recall indicate stronger physical evasion; TC-EGA is substantially stronger than the baselines.

    The mASR values for different garments and target detectors are:

    Clothing Random T-shirt Skirt Dress
    mASR 0.092 0.771 0.287 0.893

    The superscript 11 marks the YOLOv3 condition in which the input size was scaled to 50%50\% before inference. TC-EGA's mASR is approximately 1.01.0 near viewing angles 0∘0^\circ and 180∘180^\circ, but decreases near 90∘90^\circ and 270∘270^\circ because the camera captures a smaller clothing area from those directions. It outperforms the compared attacks at almost every viewing angle, whereas single AdvPatch and AdvTshirt attacks lose effectiveness as the camera rotates. Larger garments, especially dresses, are more effective because more texture is visible. The attack has comparable effectiveness indoors and outdoors but weakens as the person moves farther from the camera. The nonzero mASR values against YOLOv3, Faster R-CNN, and Mask R-CNN show transfer across detector architectures, although not perfectly.

  10. Knowl 10 — Limited cross-detector transferability

    limitation

    A texture optimized against one person detector can attack other detectors to some extent, but the paper reports that transferability is not very strong. The authors identify model ensembling as a possible way to improve cross-detector transfer, but do not evaluate that remedy as part of the presented method.

Coverage note — No substantial contributed material was deliberately omitted; background, related work, acknowledgements, and the paper's general discussion of potential negative impact were excluded as non-contributory context.

References

  1. 1.Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok. Synthesizing robust adversarial examples. In International conference on machine learning, pages 284–293. PMLR, 2018. 1, 2
  2. 2.Fred L. Bookstein. Principal warps: Thin-plate splines and the decomposition of deformations. IEEE Transactions on pattern analysis and machine intelligence, 11(6):567–585, 1989. 2
  3. 3.Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high fidelity natural image synthesis. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. 2
  4. 4.Tom B Brown, Dandelion Mane, Aurko Roy, Mart ´ ´ın Abadi, and Justin Gilmer. Adversarial patch. arXiv preprint arXiv:1712.09665, 2017. 1, 2
  5. 5.Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In 2017 IEEE Symposium on Security and Privacy (sp), pages 39–57. IEEE, 2017. 1
  6. 6.Navneet Dalal and Bill Triggs. Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), volume 1, pages 886–893. Ieee, 2005. 5
  7. 7.Gianluca Donato and Serge Belongie. Approximate thin plate spline mappings. In European conference on computer vision, pages 21–31. Springer, 2002. 2, 3
  8. 8.Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial attacks with momentum. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 9185–9193, 2018. 1
  9. 9.Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. Robust physical-world attacks on deep learning visual classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1625–1634, 2018. 1, 2
  10. 10.Ian Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. In International Conference on Learning Representations, 2015. 1, 2, 8
  11. 11.Allen Hatcher. Algebraic Topology. Cambridge University Press, 2002. 5
  12. 12.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir- ´shick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 6
  13. 13.R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio. Learning deep representations by mutual information estimation and maximization. In International Conference on Learning Representations, 2019. 3, 4
  14. 14.Yu-Chih-Tuan Hu, Bo-Han Kung, Daniel Stanley Tan, Jun-Cheng Chen, Kai-Lung Hua, and Wen-Huang Cheng. Naturalistic physical adversarial patch for object detectors. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7848–7857, 2021. 1, 2
  15. 15.Lifeng Huang, Chengying Gao, Yuyin Zhou, Cihang Xie, Alan L Yuille, Changqing Zou, and Ning Liu. Universal physical camouflage attacks on object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 720–729, 2020. 1, 2
  16. 16.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4401–4410, 2019. 2
  17. 17.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6
  18. 18.Alexey Kurakin, Ian Goodfellow, and Samy Bengio. Adversarial examples in the physical world. In International Conference on Learning Representations, 2017. 1, 2
  19. 19.Fangzhou Liao, Ming Liang, Yinpeng Dong, Tianyu Pang, Xiaolin Hu, and Jun Zhu. Defense against adversarial attacks using high-level representation guided denoiser. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1778–1787, 2018. 8
  20. 20.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 6
  21. 21.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3431–3440, 2015. 2, 4
  22. 22.Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard. Deepfool: a simple and accurate method to fool deep neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2574–2582, 2016. 1
  23. 23.Anh Nguyen, Jason Yosinski, and Jeff Clune. Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 427–436, 2015. 1
  24. 24.Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami. The limitations of deep learning in adversarial settings. In 2016 IEEE European Symposium on Security and Privacy (EuroS&P), pages 372–387. IEEE, 2016. 1
  25. 25.Alec Radford, Luke Metz, and Soumith Chintala. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434, 2015. 2
  26. 26.Joseph Redmon and Ali Farhadi. Yolo9000: better, faster, stronger. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7263–7271, 2017. 6
  27. 27.Joseph Redmon and Ali Farhadi. Yolov3: An incremental improvement. arXiv preprint arXiv:1804.02767, 2018. 6
  28. 28.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: towards real-time object detection with region proposal networks. IEEE transactions on pattern analysis and machine intelligence, 39(6):1137–1149, 2016. 6
  29. 29.Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, and Michael K Reiter. Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition. In Proceedings of the 2016 acm sigsac conference on computer and communications security, pages 1528–1540, 2016. 1, 2, 3, 4
  30. 30.Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller. Striving for simplicity: The all convolutional net. arXiv preprint arXiv:1412.6806, 2014. 2, 4
  31. 31.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In International Conference on Learning Representations, 2014. 1, 2
  32. 32.Simen Thys, Wiebe Van Ranst, and Toon Goedeme. Fooling automated surveillance cameras: adversarial patches to attack person detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 0–0, 2019. 1, 2, 3, 6, 7
  33. 33.Yi Wang, Jingyang Zhou, Tianlong Chen, Sijia Liu, Shiyu Chang, Chandrajit Bajaj, and Zhangyang Wang. Can 3d adversarial logos cloak humans? arXiv preprint arXiv:2006.14655, 2020. 2
  34. 34.Zuxuan Wu, Ser-Nam Lim, Larry S Davis, and Tom Goldstein. Making an invisibility cloak: Real world adversarial attacks on object detectors. In European Conference on Computer Vision, pages 1–17. Springer, 2020. 1, 2
  35. 35.Kaidi Xu, Gaoyuan Zhang, Sijia Liu, Quanfu Fan, Mengshu Sun, Hongge Chen, Pin-Yu Chen, Yanzhi Wang, and Xue Lin. Adversarial t-shirt! evading person detectors in a physical world. In European Conference on Computer Vision, pages 665–681. Springer, 2020. 1, 2, 3, 6, 7
  36. 36.Weilin Xu, David Evans, and Yanjun Qi. Feature squeezing: Detecting adversarial examples in deep neural networks. arXiv preprint arXiv:1704.01155, 2017. 8
  37. 37.Xiaopei Zhu, Xiao Li, Jianmin Li, Zheyao Wang, and Xiaolin Hu. Fooling thermal infrared pedestrian detectors in real world using small bulbs. In The Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI, 2021. 1

Citation

MLA
Hu, Z., et al. “Adversarial Texture for Fooling Person Detectors in the Physical World”. arXiv, 2022, http://arxiv.org/abs/2203.03373v4.
APA
Hu, Z., Huang, S., Zhu, X., Sun, F., Zhang, B., & Hu, X. (2022). Adversarial Texture for Fooling Person Detectors in the Physical World. arXiv. http://arxiv.org/abs/2203.03373v4
Chicago
Hu, Z., S. Huang, X. Zhu, F. Sun, B. Zhang, and X. Hu. 2022. “Adversarial Texture for Fooling Person Detectors in the Physical World”. arXiv. http://arxiv.org/abs/2203.03373v4.
Harvard
Hu, Z. et al. (2022) “Adversarial Texture for Fooling Person Detectors in the Physical World”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.03373v4.
Vancouver
1. Hu Z, Huang S, Zhu X, Sun F, Zhang B, Hu X (2022) Adversarial Texture for Fooling Person Detectors in the Physical World. arXiv

BibTeX

@article{hu2022adversarial,
  title = {Adversarial Texture for Fooling Person Detectors in the Physical World},
  author = {Hu, Zhanhao and Huang, Siyuan and Zhu, Xiaopei and Sun, Fuchun and Zhang, Bo and Hu, Xiaolin},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.03373v4},
  eprint = {2203.03373}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE