Image Shortcut Squeezing: Countering Perturbative Availability Poisons with Compression

Zhuoran LiuZhengyu ZhaoMartha A. Larson

article2023ICML57 citations

Proposes Image Shortcut Squeezing, a simple compression-based defense that neutralizes twelve state-of-the-art perturbative availability poisons by exploiting the frequency characteristics of poison shortcuts, matching or outperforming adversarial training with far greater efficiency.

Listen

Online image protection increasingly relies on perturbative availability poisons, which introduce imperceptible modifications to training images to prevent unauthorized machine learning models from learning useful representations. Prior literature has maintained that these data-protecting poisons are virtually impossible to neutralize without either suffering major computational overhead or incurring large penalties on model accuracy. The article challenges this widely held belief by evaluating whether simple image compression can effectively neutralize these training-time availability poisons.

To test this, the article analyzes 12 leading poisoning techniques and categorizes them by how their perturbations are generated: using slightly-trained surrogate models, fully-trained surrogate models, or surrogate-free methods. The evaluation spans multiple standard benchmarks (CIFAR-10, CIFAR-100, and a 100-class ImageNet subset) across various neural network architectures, comparing compression defenses against standard data augmentations, image filters, and robust adversarial training.

The findings show that poison patterns correlate directly with their generation method: slightly-trained models generate low-frequency color perturbations, while fully-trained models generate high-frequency patterns. Exploiting this insight, the article introduces Image Shortcut Squeezing, a defense applying simple grayscale and JPEG compression to remove these shortcut patterns during training. On CIFAR-10, this approach restores average classification accuracy to 81.73%, outperforming previous preprocessing countermeasures by an absolute margin of 37.97%. The method achieves performance comparable to adversarial training while requiring only one-seventh of the training time and generalizing across multiple perturbation bounds, including challenging one-pixel modifications where adversarial training fails completely.

These results indicate that current data-poisoning defenses for privacy and proprietary protection provide a false sense of security against model exploiters, as they can be dismantled with basic, computationally inexpensive preprocessing. In addition, tests with adaptive poisons indicate that even attack-aware modifications currently struggle to circumvent these compression defenses effectively.

For practitioners training machine learning models on potentially poisoned external data, combining grayscale and JPEG compression offers an efficient, immediate countermeasure that restores model utility with minimal accuracy loss on clean data. Researchers developing future data protection methods must incorporate compression countermeasures during benchmark evaluations. However, stakeholders should recognize that data protection is an evolving dynamic; future adaptive poisons may eventually bypass static compression rules, warranting ongoing research into automated attack identification and accuracy-preserving transformations.

No sufficiently relevant recommendations were found.

Cover for Image Shortcut Squeezing: Countering Perturbative Availability Poisons with Compression

Abstract

Perturbative availability poisons (PAPs) add small changes to images to prevent their use for model training. Current research adopts the belief that practical and effective approaches to countering PAPs do not exist. In this paper, we argue that it is time to abandon this belief. We present extensive experiments showing that 12 state-of-the-art PAP methods are vulnerable to Image Shortcut Squeezing (ISS), which is based on simple compression. For example, on average, ISS restores the CIFAR-10 model accuracy to 81.73%, surpassing the previous best preprocessing-based countermeasures by 37.97% absolute. ISS also (slightly) outperforms adversarial training and has higher generalizability to unseen perturbation norms and also higher efficiency. Our investigation reveals that the property of PAP perturbations depends on the type of surrogate model used for poison generation, and it explains why a specific ISS compression yields the best performance for a specific type of PAP perturbation. We further test stronger, adaptive poisoning, and show it falls short of being an ideal defense against ISS. Overall, our results demonstrate the importance of considering various (simple) countermeasures to ensure the meaningfulness of analysis carried out during the development of PAP methods. Our code is available at https://github.com/liuzrcc/ImageShortcutSqueezing.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Perturbative Availability Poison (PAP)
  • 2.2. PAP Countermeasures
  • 2.3. Adversarial Perturbations and Countermeasures
  • 3. Analysis of Perturbative Availability Poisons
  • 3.1. Problem Formulation
  • 3.2. Categorization of Existing PAP Methods
  • 3.3. Frequency-based Interpretation of Perturbations
  • 3.4. Our Image Shortcut Squeezing
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Evaluation in the Common Scenario
  • 4.3. Evaluation in Challenging Scenarios
  • 4.4. Adaptive Poisons to ISS
  • 4.5. Further Analyses
  • 5. Conclusion and Outlook
  • Acknowledgements
  • References
  • A. Brief Descriptions of Implemented PAP Methods
  • B. Hyperparameters for Different Countermeasures
  • C. Color Channel Difference Mitigation Methods on EM
  • D. PAP Countermeasures in Facial Recognition

Knowls

  1. Knowl 1 — Image Shortcut Squeezing matches compression to perturbation structure

    model/method

    Image Shortcut Squeezing (ISS) counters perturbative availability poisons (PAPs) by preprocessing poisoned training images before classifier training. For perturbations with low spatial frequency but large differences across color channels, ISS uses grayscale conversion: it forms a weighted sum of the color channels and copies that value to all three channels. For high-frequency perturbations, it uses JPEG compression or bit-depth reduction (BDR). The experiments use JPEG quality factor 10 and BDR at 2 bits unless specified otherwise. When the poison type is unknown, the paper also tests applying grayscale and JPEG together (Gray+JPEG) as a broadly effective option. ISS changes the training images, not the target model architecture.

  2. Knowl 2 — Poison frequency patterns correspond to surrogate training stage

    empirical result

    The paper groups 12 PAP methods by the model used to generate their perturbations. Slightly-trained-surrogate methods are DC, NTGA, EM, REM, and SG; fully-trained-surrogate methods are TC, HYPO, TAP, and SEP; surrogate-free methods are LSP, AR, and OPS. In the CIFAR-10 visualizations, slightly-trained-surrogate methods tend to produce low-spatial-frequency perturbations with large color-channel differences, whereas fully-trained-surrogate methods tend to produce high-frequency perturbations. Experiments generating perturbations at different surrogate training epochs show increasing spatial frequency at later epochs for both error-minimizing and error-maximizing objectives. The paper interprets this pattern as consistent with neural networks fitting lower frequencies earlier in training; it is an empirical association, not a claim that every poison follows the pattern. Among surrogate-free methods, LSP uses upsampled Gaussian patterns and is low-frequency, AR produces texture-like high-frequency patterns, and OPS perturbs a single pixel.

  3. Knowl 3 — ISS substantially restores CIFAR-10 accuracy across PAP methods

    empirical result

    In the common setting—CIFAR-10, a ResNet-18 surrogate and target, the full training set poisoned, and L∞=8L_\infty=8—the paper reports that ISS restores clean-test accuracy to 81.73% on average across 12 PAP methods. This is 37.97 percentage points above the strongest previously studied preprocessing-based countermeasures on average. The gains depend on choosing a compression suited to the poison: grayscale gives 93.07% for DC, 93.01% for EM, and 92.84% for REM, while JPEG gives 83.87% for TAP, 84.37% for SEP, and 85.15% for AR. Gray+JPEG also restores accuracy across all 12 methods, with reported results from 69.86% to 83.79%. The clean, unpoisoned ResNet-18 accuracy is 94.68%; with Gray+JPEG it is 83.79%.

  4. Knowl 4 — ISS is faster than adversarial training and handles multiple poison norms

    empirical result

    On CIFAR-10, the paper finds adversarial training (AT) comparable to ISS for poisons constrained by L∞L_\infty and L2L_2, but substantially weaker for the tested L0L_0 poison OPS: AT yields 14.41% clean-test accuracy, while median filtering yields 85.16%. Compression-based ISS is therefore less tied to the particular perturbation norm than AT in these experiments, although no single compression operation is best for every poison. ISS training takes one-seventh as long as AT on CIFAR-10. AT in the comparison uses PGD-10 with step size 2/2552/255 and 100 training epochs.

  5. Knowl 5 — Adaptive poisons can bypass individual ISS operations but are not uniformly effective

    empirical result

    The paper tests adaptive EM poisons for L∞=8L_\infty=8 by including differentiable grayscale and/or JPEG operations in poison optimization, and adaptive LSP poisons by using 16×1616\times16 patches with equal values across color channels. Adaptive EM-Grayscale reduces target accuracy to 16.60% when grayscale is applied, showing that it can bypass that operation; however, its accuracy is 76.71% under JPEG and 74.16% under Gray+JPEG. Adaptive EM-JPEG yields 83.11% under JPEG, so it does not effectively bypass JPEG in this experiment. For EM-JPEG evaluated against BPDA with JPEG quality factor 10, the paper reports 83.70% accuracy. Adaptive LSP-G&J produces 93.01% accuracy even without ISS, indicating that its additional constraints have already weakened its poisoning effect; Gray+JPEG yields 82.13%. The authors caution that better-designed adaptive poisons may bypass ISS in the future.

  6. Knowl 6 — Compression improves poisoned-model accuracy on unseen target architectures

    empirical result

    In the CIFAR-10 transfer setting, poisons are generated without matching the target architecture, and target models include ResNet-34, VGG-19, DenseNet-121, MobileNet-V2, and ViT. Compression improves accuracy for many poison–architecture combinations, though the amount of recovery varies and is often lower for ViT than for the CNN targets. For example, EM-poisoned models have unprocessed accuracies of 18.84–34.70% across these five architectures; grayscale raises the CNN results to 82.81–87.03% and the ViT result to 63.28%. The paper finds no clear relationship between transfer difficulty and architectural similarity: transfer from ResNet-18 to ViT is not consistently harder than transfer to another CNN.

  7. Knowl 7 — ISS remains useful in partial-poisoning settings

    empirical result

    The paper evaluates random partial poisoning and poisoning all training samples from one CIFAR-10 class. With random poisoning, PAPs generally reduce accuracy substantially only when a large fraction of training data is poisoned; the authors report that even at 80% poisoning the average accuracy reduction is about 10 percentage points. Applying ISS improves accuracy in this setting across the tested methods. In the class-specific automobile experiment, unprocessed accuracy is 0.10% for EM, 0.00% for TAP, and 0.00% for SEP; JPEG preprocessing raises these results to 94.30%, 93.90%, and 94.70%, respectively. The results also show that performance depends on matching the compression to the poison: grayscale alone does not restore these three class-specific cases.

  8. Knowl 8 — ISS effects extend to CIFAR-100, an ImageNet subset, and larger perturbation bounds

    empirical result

    The paper reports ISS results beyond the main CIFAR-10 setting, with the best operation varying by poison. On CIFAR-100, EM accuracy rises from 7.25% without compression to 67.46% with grayscale; TAP rises from 9.00% to 83.77% with JPEG. On the evaluated ImageNet subset, EM rises from 31.52% to 49.78% with grayscale, and REM rises from 11.12% to 44.70% with grayscale. The ImageNet subset uses 20% of the training images from the first 100 ImageNet classes and their corresponding validation images; DC and NTGA are tested on only two classes. For the larger CIFAR-10 bound L∞=16L_\infty=16, EM rises from 16.33% without compression to 60.85% with grayscale or 63.44% with JPEG, while TAP rises from 10.98% to 84.19% with JPEG. These results support transfer across datasets and bounds, but not uniform restoration to clean accuracy.

  9. Knowl 9 — Poisoned-image testing supports shortcut suppression as ISS's mechanism

    empirical result

    To test whether PAP perturbations are learned as predictive shortcuts, the authors train and test models on poisoned images. Across the 12 methods, test accuracy on poisoned images is 93.79–100.00%, despite low accuracy on clean test images for most of the same models. When ISS is applied to the poisoned test images, accuracy falls to 10.06–25.39%. This paired evaluation supports the interpretation that models learned predictive information from the poison and that preprocessing can suppress that information at inference time.

  10. Knowl 10 — Mixed-poison experiments reveal unequal shortcut preference

    empirical result

    The paper examines low-frequency EM and high-frequency TAP perturbations together. When a model is trained on perturbations formed by adding the two methods' perturbations and is then tested separately on EM- and TAP-poisoned images, it reaches high accuracy on EM but not TAP; the authors interpret this as a preference for the EM shortcut during training. In a separate experiment, the authors average EM and TAP perturbations to create a mixed poison. On CIFAR-10, accuracy is 36.07% without preprocessing, 18.93% with grayscale, and 84.62% with JPEG. Thus, the tested mixed poison remains counterable, and the preferred compression depends on how the perturbations are combined.

Coverage note — The supplementary comparison of alternative color-channel reducers and the check comparing training-only compression with compression applied to both training and test images are omitted because they are secondary robustness analyses rather than separate central findings.

References

  1. 1.Athalye, A., Carlini, N., and Wagner, D. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In ICML, 2018a.
  2. 2.Athalye, A., Engstrom, L., Ilyas, A., and Kwok, K. Synthesizing robust adversarial examples. In ICML, 2018b.
  3. 3.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. In NeurIPS, 2020.
  4. 4.Carlini, N., Athalye, A., Papernot, N., Brendel, W., Rauber, J., Tsipras, D., Goodfellow, I., Madry, A., and Kurakin, A. On evaluating adversarial robustness. arXiv:1902.06705, 2019.
  5. 5.Chen, S., Yuan, G., Cheng, X., Gong, Y., Qin, M., Wang, Y., and Huang, X. Self-ensemble protection: Training checkpoints are good data protectors. In ICLR, 2023.
  6. 6.Cherepanova, V., Goldblum, M., Foley, H., Duan, S., Dickerson, J., Taylor, G., and Goldstein, T. Lowkey: Leveraging adversarial attacks to protect social media users from facial recognition. In ICLR, 2021.
  7. 7.Das, N., Shanbhogue, M., Chen, S.-T., Hohman, F., Chen, L., Kounavis, M. E., and Chau, D. H. Keeping the bad guys out: Protecting and vaccinating deep learning with jpeg compression. arXiv:1705.02900, 2017.
  8. 8.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In CVPR, pp. 248–255, 2009.
  9. 9.Dodge, S. and Karam, L. Understanding how image quality affects deep neural networks. In QoMEX, 2016.
  10. 10.Dong, Y., Pang, T., Su, H., and Zhu, J. Evading defenses to transferable adversarial examples by translation-invariant attacks. In CVPR, 2019.
  11. 11.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  12. 12.Dziugaite, G. K., Ghahramani, Z., and Roy, D. M. A study of the effect of jpg compression on adversarial images. arXiv:1608.00853, 2016.
  13. 13.Evtimov, I., Covert, I., Kusupati, A., and Kohno, T. Disrupting model training with adversarial shortcuts. In ICML Workshop AML, 2021.
  14. 14.Feng, J., Cai, Q.-Z., and Zhou, Z.-H. Learning to confuse: generating training time adversarial data with auto-encoder. In NeurIPS, 2019.
  15. 15.Fowl, L., Chiang, P.-y., Goldblum, M., Geiping, J., Bansal, A., Czaja, W., and Goldstein, T. Preventing unauthorized use of proprietary data: Poisoning for secure dataset release. arXiv:2103.02683, 2021a.
  16. 16.Fowl, L., Goldblum, M., Chiang, P.-y., Geiping, J., Czaja, W., and Goldstein, T. Adversarial examples make strong poisons. In NeurIPS, 2021b.
  17. 17.Fu, S., He, F., Liu, Y., Shen, L., and Tao, D. Robust unlearnable examples: Protecting data privacy against adversarial learning. In ICLR, 2021.
  18. 18.Guo, C., Frank, J. S., and Weinberger, K. Q. Low frequency adversarial perturbation. In UAI, 2019.
  19. 19.He, H., Zha, K., and Katabi, D. Indiscriminate poisoning attacks on unsupervised contrastive learning. In ICLR, 2023.
  20. 20.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, 2016.
  21. 21.Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. Densely connected convolutional networks. In CVPR, 2017.
  22. 22.Huang, H., Ma, X., Erfani, S. M., Bailey, J., and Wang, Y. Unlearnable examples: Making personal data unexploitable. In ICLR, 2021.
  23. 23.Huang, W. R., Geiping, J., Fowl, L., Taylor, G., and Goldstein, T. MetaPoison: Practical general-purpose clean-label data poisoning. In NeurIPS, 2020.
  24. 24.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. In NeurIPS, 2018.
  25. 25.Jaderberg, M., Simonyan, K., Zisserman, A., et al. Spatial transformer networks. In NeurIPS, 2015.
  26. 26.Krizhevsky, A. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.
  27. 27.Laidlaw, C., Singla, S., and Feizi, S. Perceptual adversarial robustness: Defense against unseen threat models. In ICLR, 2021.
  28. 28.Larson, M., Liu, Z., Brugman, S., and Zhao, Z. Pixel privacy. increasing image appeal while blocking automatic inference of sensitive scene information. In MediaEval, 2018.
  29. 29.Feng, J., Cai, Q.-Z., and Zhou, Z.-H. Learning to confuse: generating training time adversarial data with auto-encoder. In NeurIPS, 2019.
  30. 30.Fowl, L., Chiang, P.-y., Goldblum, M., Geiping, J., Bansal, A., Czaja, W., and Goldstein, T. Preventing unauthorized use of proprietary data: Poisoning for secure dataset release. arXiv:2103.02683, 2021a.
  31. 31.Fowl, L., Goldblum, M., Chiang, P.-y., Geiping, J., Czaja, W., and Goldstein, T. Adversarial examples make strong poisons. In NeurIPS, 2021b.
  32. 32.Fu, S., He, F., Liu, Y., Shen, L., and Tao, D. Robust unlearnable examples: Protecting data privacy against adversarial learning. In ICLR, 2021.
  33. 33.Guo, C., Frank, J. S., and Weinberger, K. Q. Low frequency adversarial perturbation. In UAI, 2019.
  34. 34.He, H., Zha, K., and Katabi, D. Indiscriminate poisoning attacks on unsupervised contrastive learning. In ICLR, 2023.
  35. 35.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In CVPR, 2016.
  36. 36.Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q. Densely connected convolutional networks. In CVPR, 2017.
  37. 37.Huang, H., Ma, X., Erfani, S. M., Bailey, J., and Wang, Y. Unlearnable examples: Making personal data unexploitable. In ICLR, 2021.
  38. 38.Huang, W. R., Geiping, J., Fowl, L., Taylor, G., and Goldstein, T. MetaPoison: Practical general-purpose clean-label data poisoning. In NeurIPS, 2020.
  39. 39.Jacot, A., Gabriel, F., and Hongler, C. Neural tangent kernel: Convergence and generalization in neural networks. In NeurIPS, 2018.
  40. 40.Jaderberg, M., Simonyan, K., Zisserman, A., et al. Spatial transformer networks. In NeurIPS, 2015.
  41. 41.Krizhevsky, A. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009.
  42. 42.Laidlaw, C., Singla, S., and Feizi, S. Perceptual adversarial robustness: Defense against unseen threat models. In ICLR, 2021.
  43. 43.Larson, M., Liu, Z., Brugman, S., and Zhao, Z. Pixel privacy. increasing image appeal while blocking automatic inference of sensitive scene information. In MediaEval, 2018.
  44. 44.LeCun, Y., Bengio, Y., and Hinton, G. Deep learning. Nature, 521(7553):436–444, 2015.
  45. 45.Li, C. Y., Shamsabadi, A. S., Sanchez-Matilla, R., Mazzon, R., and Cavallaro, A. Scene privacy protection. In ICASSP, 2019.
  46. 46.Liu, Z., Zhao, Z., Larson, M., and Amsaleg, L. Exploring quality camouflage for social images. In MediaEval, 2020.
  47. 47.Luo, T., Ma, Z., Xu, Z.-Q. J., and Zhang, Y. Theory of the frequency principle for general deep neural networks. CSIAM Trans. Appl. Math., 2021.
  48. 48.Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. Towards deep learning models resistant to adversarial attacks. In ICLR, 2018.
  49. 49.Naseer, M. M., Khan, S. H., Khan, M. H., Shahbaz Khan, F., and Porikli, F. Cross-domain transferability of adversarial perturbations. In NeurIPS, 2019.
  50. 50.Oh, S. J., Benenson, R., Fritz, M., and Schiele, B. Faceless person recognition: Privacy implications in social media. In ECCV, 2016.
  51. 51.Oh, S. J., Fritz, M., and Schiele, B. Adversarial image perturbation for privacy protection a game theory perspective. In ICCV, 2017.
  52. 52.Orekondy, T., Schiele, B., and Fritz, M. Towards a visual privacy advisor: Understanding and predicting privacy risks in images. In ICCV, 2017.
  53. 53.Radiya-Dixit, E., Hong, S., Carlini, N., and Tramer, F. Data poisoning won’t save you from facial recognition. In ICLR, 2022.
  54. 54.Rahaman, N., Baratin, A., Arpit, D., Draxler, F., Lin, M., Hamprecht, F., Bengio, Y., and Courville, A. On the spectral bias of neural networks. In ICML, 2019.
  55. 55.Rajabi, A., Bobba, R. B., Rosulek, M., Wright, C., and Feng, W.-c. On the (im) practicality of adversarial perturbation for image privacy. In PETS, 2021.
  56. 56.Ren, J., Xu, H., Wan, Y., Ma, X., Sun, L., and Tang, J. Transferable unlearnable examples. In ICLR, 2023.
  57. 57.Ronneberger, O., Fischer, P., and Brox, T. U-net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015.
  58. 58.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, 2018.
  59. 59.Sandoval-Segura, P., Singla, V., Geiping, J., Goldblum, M., Goldstein, T., and Jacobs, D. W. Autoregressive perturbations for data poisoning. In NeurIPS, 2022.
  60. 60.Sattar, H., Krombholz, K., Pons-Moll, G., and Fritz, M. Body shape privacy in images: Understanding privacy and preventing automatic shape extraction. In ECCV, 2020.
  61. 61.Schmidhuber, J. Deep learning in neural networks: An overview. Neural Networks, 61:85–117, 2015.
  62. 62.Shah, H., Tamuly, K., Raghunathan, A., Jain, P., and Netrapalli, P. The pitfalls of simplicity bias in neural networks. In NeurIPSw, 2020.
  63. 63.Shan, S., Wenger, E., Zhang, J., Li, H., Zheng, H., and Zhao, B. Y. Fawkes: Protecting privacy against unauthorized deep learning models. In USENIX Security, 2020.
  64. 64.Shen, J., Zhu, X., and Ma, D. TensorClog: An imperceptible poisoning attack on deep neural network applications. IEEE Access, 7:41498–41506, 2019.
  65. 65.Shin, R. and Song, D. Jpeg-resistant adversarial images. In NeurIPSw, 2017.
  66. 66.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015.
  67. 67.Tao, L., Feng, L., Yi, J., Huang, S.-J., and Chen, S. Better safe than sorry: Preventing delusive adversaries with adversarial training. In NeurIPS, 2021.
  68. 68.Tian, S., Yang, G., and Cai, Y. Detecting adversarial examples through image transformation. In AAAI, 2018.
  69. 69.Tramer, F. and Boneh, D. Adversarial training and robustness for multiple perturbations. In NeurIPS, 2019.
  70. 70.Tramer, F., Carlini, N., Brendel, W., and Madry, A. On adaptive attacks to adversarial example defenses. In NeurIPS, 2020.
  71. 71.van Vlijmen, D., Kolmus, A., Liu, Z., Zhao, Z., and Larson, M. Generative poisoning using random discriminators. In ECCVw, 2022.
  72. 72.Wang, Z., Wang, Y., and Wang, Y. Fooling adversarial training with inducing noise. arXiv:2111.10130, 2021.
  73. 73.Wen, R., Zhao, Z., Liu, Z., Backes, M., Wang, T., and Zhang, Y. Is adversarial training really a silver bullet for mitigating data poisoning? In ICLR, 2023.
  74. 74.Wu, S., Chen, S., Xie, C., and Huang, X. One-pixel shortcut: on the learning preference of deep neural networks. In ICLR, 2023.
  75. 75.Xie, C., Wang, J., Zhang, Z., Ren, Z., and Yuille, A. Mitigating adversarial effects through randomization. ICLR, 2018.
  76. 76.Xie, Y. and Richmond, D. Pre-training on grayscale imagenet improves medical image classification. In ECCVw, 2018.
  77. 77.Xu, W., Evans, D., and Qi, Y. Feature squeezing: Detecting adversarial examples in deep neural networks. In NDSS, 2017.
  78. 78.Xu, Z.-Q. J., Zhang, Y., and Xiao, Y. Training behavior of deep neural network in frequency domain. In ICONIP, 2019.
  79. 79.Yu, D., Zhang, H., Chen, W., Yin, J., and Liu, T.-Y. Availability attacks create shortcuts. In KDD, 2022.
  80. 80.Yuan, C.-H. and Wu, S.-H. Neural tangent generalization attacks. In ICML, 2021.
  81. 81.Zhang, H., Yu, Y., Jiao, J., Xing, E., El Ghaoui, L., and Jordan, M. Theoretically principled trade-off between robustness and accuracy. In ICML, 2019.
  82. 82.Zhang, J., Ma, X., Yi, Q., Sang, J., Jiang, Y., Wang, Y., and Xu, C. Unlearnable clusters: Towards label-agnostic unlearnable examples. In CVPR, 2023.

Citation

MLA
Liu, Z., et al. “Image Shortcut Squeezing: Countering Perturbative Availability Poisons with Compression”. International Conference on Machine Learning, vol. 202, 2023, pp. 22473–87, https://proceedings.mlr.press/v202/liu23bb.html.
APA
Liu, Z., Zhao, Z., & Larson, M. (2023). Image Shortcut Squeezing: Countering Perturbative Availability Poisons with Compression. International Conference on Machine Learning, 202, 22473–22487. https://proceedings.mlr.press/v202/liu23bb.html
Chicago
Liu, Z., Z. Zhao, and M. Larson. 2023. “Image Shortcut Squeezing: Countering Perturbative Availability Poisons with Compression”. International Conference on Machine Learning 202: 22473–87. https://proceedings.mlr.press/v202/liu23bb.html.
Harvard
Liu, Z., Zhao, Z. and Larson, M. (2023) “Image Shortcut Squeezing: Countering Perturbative Availability Poisons with Compression”, International Conference on Machine Learning. PMLR, pp. 22473–22487. Available at: https://proceedings.mlr.press/v202/liu23bb.html.
Vancouver
1. Liu Z, Zhao Z, Larson M (2023) Image Shortcut Squeezing: Countering Perturbative Availability Poisons with Compression. In: International Conference on Machine Learning. PMLR, pp 22473–22487

BibTeX

@InProceedings{pmlr-v202-liu23bb,
  title = 	 {Image Shortcut Squeezing: Countering Perturbative Availability Poisons with Compression},
  author =       {Liu, Zhuoran and Zhao, Zhengyu and Larson, Martha},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {22473--22487},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/liu23bb/liu23bb.pdf},
  url = 	 {https://proceedings.mlr.press/v202/liu23bb.html},
  abstract = 	 {Perturbative availability poisoning (PAP) adds small changes to images to prevent their use for model training. Current research adopts the belief that practical and effective approaches to countering such poisons do not exist. In this paper, we argue that it is time to abandon this belief. We present extensive experiments showing that 12 state-of-the-art PAP methods are vulnerable to Image Shortcut Squeezing (ISS), which is based on simple compression. For example, on average, ISS restores the CIFAR-10 model accuracy to 81.73%, surpassing the previous best preprocessing-based countermeasures by 37.97% absolute. ISS also (slightly) outperforms adversarial training and has higher generalizability to unseen perturbation norms and also higher efficiency. Our investigation reveals that the property of PAP perturbations depends on the type of surrogate model used for poison generation, and it explains why a specific ISS compression yields the best performance for a specific type of PAP perturbation. We further test stronger, adaptive poisoning, and show it falls short of being an ideal defense against ISS. Overall, our results demonstrate the importance of considering various (simple) countermeasures to ensure the meaningfulness of analysis carried out during the development of availability poisons.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/